Description
NVIDIA's Deep Learning Software team is looking for a Senior Technical Program Manager to lead programs across model pre-training, production RL runs, evaluation, and agentic AI infrastructure.
You will lead multi-functional programs across training frameworks, evaluation environments, agent and model runtimes, datasets, verifiers, and distributed training infrastructure.
Key responsibilities include:
- Partnering with AI researchers, engineering leaders, product, infrastructure, and QA teams to define roadmaps, achievements, release plans, and measurable success criteria.
- Coordinating large-scale RL training and evaluation experiments, including handling GPU resources, dependency tracking, run scheduling, results reporting, release readiness, technical decisions, integration plans, and program updates.
Requirements:
- Bachelor's degree in computer science, engineering, or a related technical field, or equivalent experience.
- 10+ years of technical program management, engineering program management, or related experience delivering sophisticated software platforms.
- Experience leading global, matrixed programs across research, software engineering, infrastructure, QA, release teams, and partner groups.
- Strong understanding of the AI model lifecycle, including training, post-training, evaluation, experimentation, production readiness, reinforcement learning concepts, and GPU-accelerated distributed systems.
- Experience running software releases across repositories, dependencies, test configurations, quality gates, collaborator approvals, open-source workflows, CI/CD systems, and tools such as GitHub, Git, Jira, Linear, Aha!, or Confluence.
Nice to have:
- Experience supporting reinforcement learning, post-training, agentic AI, or large-scale model evaluation programs.
- Familiarity with PPO, GRPO, asynchronous RL, distributed inference, rollout generation, policy optimization, evaluation harnesses, verifiers, benchmark development, or reproducible experimentation.
- Knowledge of GPU infrastructure, distributed training, Kubernetes, workload schedulers, cluster capacity management, performance analysis, open-source contributor workflows, release readiness, or operational metrics.
You will also be eligible for equity and benefits.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Technical-Program-Manager--GenAI-and-Models_JR2022028