# AI Infrastructure Software Engineer — CosmosLab

**Company**: NVIDIA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Beijing/AI-Infrastructure-Software-Engineer---CosmosLab_JR2020792?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_5a99560c-428

## Description

NVIDIA is seeking an AI Infrastructure Software Engineer to join the Cosmos Lab Infra team. The successful candidate will design, assemble, and improve the infrastructure for large-scale AI training, spanning pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL) post-training.

**Responsibilities:**

- Create and implement the training infrastructure spanning pre-training, SFT, and RL post-training for Physical AI world foundation models.

- Develop and improve the pre-training and SFT pipelines , large-scale data loading, distributed training, and checkpointing , to achieve high throughput and scalability.

- Develop and improve the inference and evaluation stack, including the inference engine, inference/generation pipelines, and evaluation pipelines.

- Build and improve the effective interaction and data flow among the RL system's roles.

- Integrate and orchestrate simulation and robotics environments as RL environments.

- Build and refine the distributed training backend.

- Improve the efficiency, scalability, and resiliency of training and RL workloads.

- Define meaningful, actionable reliability and efficiency metrics to track and improve system reliability.

- Root cause, triage, and resolve failures from the application level down to the framework, GPU, and network/hardware level.

**Requirements:**

- 5+ years developing software infrastructure for large-scale AI or distributed systems.

- Bachelor's degree or higher in Computer Science or a related technical field.

- Strong debugging and triage skills across the stack.

- Proven track record building and scaling large-scale distributed systems.

- Hands-on experience with AI training and/or inference infrastructure.

- Proficiency in Python and solid software engineering practices.

- Excellent communication and collaboration skills.

**Nice to Have:**

- Experience building RL / post-training infrastructure.

- Background with building large-scale, production-grade pre-training / SFT infrastructure.

- Experience integrating simulation / robotics environments into training or RL loops.

- Comprehensive knowledge of DL framework internals.

- Proficiency in C/C++/CUDA for performance-critical components and custom kernels.

NVIDIA offers highly competitive salaries and a comprehensive benefits package.

## Skills

### Required
- Python
- distributed systems
- AI training
- inference infrastructure
- debugging
- software engineering

### Nice to have
- RL infrastructure
- pre-training infrastructure
- simulation environments
- DL framework internals
- C/C++/CUDA

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Beijing/AI-Infrastructure-Software-Engineer---CosmosLab_JR2020792?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
