Description
NVIDIA is seeking an AI Infrastructure Software Engineer to join the Cosmos Lab Infra team. The successful candidate will design, assemble, and improve the infrastructure for large-scale AI training, spanning pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL) post-training.
Responsibilities:
- Create and implement the training infrastructure spanning pre-training, SFT, and RL post-training for Physical AI world foundation models.
- Develop and improve the pre-training and SFT pipelines , large-scale data loading, distributed training, and checkpointing , to achieve high throughput and scalability.
- Develop and improve the inference and evaluation stack, including the inference engine, inference/generation pipelines, and evaluation pipelines.
- Build and improve the effective interaction and data flow among the RL system's roles.
- Integrate and orchestrate simulation and robotics environments as RL environments.
- Build and refine the distributed training backend.
- Improve the efficiency, scalability, and resiliency of training and RL workloads.
- Define meaningful, actionable reliability and efficiency metrics to track and improve system reliability.
- Root cause, triage, and resolve failures from the application level down to the framework, GPU, and network/hardware level.
Requirements:
- 5+ years developing software infrastructure for large-scale AI or distributed systems.
- Bachelor's degree or higher in Computer Science or a related technical field.
- Strong debugging and triage skills across the stack.
- Proven track record building and scaling large-scale distributed systems.
- Hands-on experience with AI training and/or inference infrastructure.
- Proficiency in Python and solid software engineering practices.
- Excellent communication and collaboration skills.
Nice to Have:
- Experience building RL / post-training infrastructure.
- Background with building large-scale, production-grade pre-training / SFT infrastructure.
- Experience integrating simulation / robotics environments into training or RL loops.
- Comprehensive knowledge of DL framework internals.
- Proficiency in C/C++/CUDA for performance-critical components and custom kernels.
NVIDIA offers highly competitive salaries and a comprehensive benefits package.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Beijing/AI-Infrastructure-Software-Engineer---CosmosLab_JR2020792