New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

NVIDIA Intern - Cosmos Lab Infrastructure

NVIDIA
Apply →
internship internship

First indexed 18 Sept 2026

Description

Join NVIDIA's Cosmos Lab Infrastructure team as an intern to develop training and post-training systems for advanced Physical AI models. You will work on a focused project, implementing and evaluating systems improvements on real AI workloads using NVIDIA's GPU infrastructure.

Responsibilities:

  • Develop and optimize training infrastructure for advanced Physical AI world models, supporting pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL).
  • Build Physical AI post-training and RL infrastructure supporting advanced training algorithms.
  • Improve efficiency and scalability across training, inference, simulation, and evaluation through scheduling, placement, dynamic resource allocation, and load balancing.
  • Analyze and optimize system performance, working with researchers to investigate, support, and compare emerging Physical AI models, training workflows, and algorithms from a systems perspective.

Requirements:

  • Pursuing a Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • Strong Python and debugging skills, with systems fundamentals in concurrency, distributed execution, memory management, or data movement.
  • Practical experience in at least one area: training infrastructure, RL infrastructure, simulation or robotics integration, or inference infrastructure.
  • Strong analytical and communication skills, curiosity, and a willingness to learn.

Nice to Have:

  • Experience optimizing training infrastructure, including distributed parallelism, low-precision training, GPU memory efficiency, or compute-communication overlap.
  • Experience optimizing scheduling, placement, resource allocation, or data transfer across training, rollout, simulation, and evaluation.
  • Experience extending RL pipelines, integrating simulation environments or robot interfaces, or optimizing inference; GPU profiling, C++/CUDA development, and open-source contributions or research in ML systems.