New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Software Engineer, RL Post-Training Frameworks

NVIDIA
Apply →
remote senior full-time

First indexed 3 Jul 2026

Description

Reinforcement learning post-training is driving significant capability gains in AI today. It teaches a model to reason through hard problems, follow complex instructions, and act autonomously. NVIDIA is building an RL Frameworks engineering team to develop open-source tools and infrastructure for AI researchers and post-training teams.

You will architect and build RL post-training infrastructure that scales efficiently from experimentation on a single GPU to production across thousands of nodes. This involves tuning RL training-inference-rollout loops on GPUs, CPUs, and LPUs for performance, contributing to and improving open-source RL frameworks, and partnering with teams who own them.

The role also spans fault tolerance, elastic scaling, and fast restarts for long-running distributed training jobs. You will work with researchers to understand and address their needs, optimize deep learning frameworks, and build distributed infrastructure.

Responsibilities:

  • Architect and build scalable RL post-training infrastructure
  • Tune RL training-inference-rollout loops on GPUs, CPUs, and LPUs
  • Contribute to and improve open-source RL frameworks
  • Partner with teams who own RL frameworks
  • Ensure fault tolerance, elastic scaling, and fast restarts for distributed training jobs

Requirements:

  • MS or PhD in Computer Science, Computer Engineering, or a related field
  • 5+ years of professional experience in distributed systems, high-performance computing, deep learning infrastructure, or ML systems engineering
  • Strong proficiency in Python and C/C++
  • Demonstrated experience building or contributing to large-scale distributed systems or runtime frameworks
  • Strong verbal and written communication skills

Nice to Have:

  • Experience with reinforcement learning for LLM post-training
  • Knowledge of PyTorch internals and distributed training primitives
  • Familiarity with Kubernetes runtime internals
  • Experience with end-to-end distributed systems design

Benefits: NVIDIA offers highly competitive salaries and a comprehensive benefits package. Learn more at www.nvidiabenefits.com/