# NVIDIA Intern - Cosmos Lab Infrastructure

**Company**: NVIDIA
**Experience**: internship
**Job type**: internship
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Beijing/AI-Infrastructure-and-Frameworks-Intern--Cosmos-Lab---2027_JR2025559?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_1bf61cde-c69

## Description

Join NVIDIA's Cosmos Lab Infrastructure team as an intern to develop training and post-training systems for advanced Physical AI models. You will work on a focused project, implementing and evaluating systems improvements on real AI workloads using NVIDIA's GPU infrastructure.

**Responsibilities:**

- Develop and optimize training infrastructure for advanced Physical AI world models, supporting pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL).

- Build Physical AI post-training and RL infrastructure supporting advanced training algorithms.

- Improve efficiency and scalability across training, inference, simulation, and evaluation through scheduling, placement, dynamic resource allocation, and load balancing.

- Analyze and optimize system performance, working with researchers to investigate, support, and compare emerging Physical AI models, training workflows, and algorithms from a systems perspective.

**Requirements:**

- Pursuing a Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field.

- Strong Python and debugging skills, with systems fundamentals in concurrency, distributed execution, memory management, or data movement.

- Practical experience in at least one area: training infrastructure, RL infrastructure, simulation or robotics integration, or inference infrastructure.

- Strong analytical and communication skills, curiosity, and a willingness to learn.

**Nice to Have:**

- Experience optimizing training infrastructure, including distributed parallelism, low-precision training, GPU memory efficiency, or compute-communication overlap.

- Experience optimizing scheduling, placement, resource allocation, or data transfer across training, rollout, simulation, and evaluation.

- Experience extending RL pipelines, integrating simulation environments or robot interfaces, or optimizing inference; GPU profiling, C++/CUDA development, and open-source contributions or research in ML systems.

## Skills

### Required
- Python
- debugging
- concurrency
- distributed execution
- memory management
- data movement
- Computer Science
- Computer Engineering
- Electrical Engineering

### Nice to have
- distributed parallelism
- low-precision training
- GPU memory efficiency
- compute-communication overlap
- scheduling
- placement
- resource allocation
- data transfer
- reinforcement learning
- simulation environments
- robot interfaces
- inference
- GPU profiling
- C++/CUDA development

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Beijing/AI-Infrastructure-and-Frameworks-Intern--Cosmos-Lab---2027_JR2025559?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
