New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

AI Infrastructure Software Engineer — CosmosLab

NVIDIA
Apply →
senior full-time

First indexed 8 Jul 2026

Description

NVIDIA is seeking an AI Infrastructure Software Engineer to join the Cosmos Lab Infra team. The successful candidate will design, assemble, and improve the infrastructure for large-scale AI training, spanning pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL) post-training.

Responsibilities:

  • Create and implement the training infrastructure spanning pre-training, SFT, and RL post-training for Physical AI world foundation models.
  • Develop and improve the pre-training and SFT pipelines , large-scale data loading, distributed training, and checkpointing , to achieve high throughput and scalability.
  • Develop and improve the inference and evaluation stack, including the inference engine, inference/generation pipelines, and evaluation pipelines.
  • Build and improve the effective interaction and data flow among the RL system's roles.
  • Integrate and orchestrate simulation and robotics environments as RL environments.
  • Build and refine the distributed training backend.
  • Improve the efficiency, scalability, and resiliency of training and RL workloads.
  • Define meaningful, actionable reliability and efficiency metrics to track and improve system reliability.
  • Root cause, triage, and resolve failures from the application level down to the framework, GPU, and network/hardware level.

Requirements:

  • 5+ years developing software infrastructure for large-scale AI or distributed systems.
  • Bachelor's degree or higher in Computer Science or a related technical field.
  • Strong debugging and triage skills across the stack.
  • Proven track record building and scaling large-scale distributed systems.
  • Hands-on experience with AI training and/or inference infrastructure.
  • Proficiency in Python and solid software engineering practices.
  • Excellent communication and collaboration skills.

Nice to Have:

  • Experience building RL / post-training infrastructure.
  • Background with building large-scale, production-grade pre-training / SFT infrastructure.
  • Experience integrating simulation / robotics environments into training or RL loops.
  • Comprehensive knowledge of DL framework internals.
  • Proficiency in C/C++/CUDA for performance-critical components and custom kernels.

NVIDIA offers highly competitive salaries and a comprehensive benefits package.