New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Technical Program Manager, GenAI and Models

NVIDIA
Apply →
remote senior full-time Santa Clara, CA

First indexed 29 Jul 2026

Description

NVIDIA's Deep Learning Software team is looking for a Senior Technical Program Manager to lead programs across model pre-training, production RL runs, evaluation, and agentic AI infrastructure.

You will lead multi-functional programs across training frameworks, evaluation environments, agent and model runtimes, datasets, verifiers, and distributed training infrastructure.

Key responsibilities include:

  • Partnering with AI researchers, engineering leaders, product, infrastructure, and QA teams to define roadmaps, achievements, release plans, and measurable success criteria.
  • Coordinating large-scale RL training and evaluation experiments, including handling GPU resources, dependency tracking, run scheduling, results reporting, release readiness, technical decisions, integration plans, and program updates.

Requirements:

  • Bachelor's degree in computer science, engineering, or a related technical field, or equivalent experience.
  • 10+ years of technical program management, engineering program management, or related experience delivering sophisticated software platforms.
  • Experience leading global, matrixed programs across research, software engineering, infrastructure, QA, release teams, and partner groups.
  • Strong understanding of the AI model lifecycle, including training, post-training, evaluation, experimentation, production readiness, reinforcement learning concepts, and GPU-accelerated distributed systems.
  • Experience running software releases across repositories, dependencies, test configurations, quality gates, collaborator approvals, open-source workflows, CI/CD systems, and tools such as GitHub, Git, Jira, Linear, Aha!, or Confluence.

Nice to have:

  • Experience supporting reinforcement learning, post-training, agentic AI, or large-scale model evaluation programs.
  • Familiarity with PPO, GRPO, asynchronous RL, distributed inference, rollout generation, policy optimization, evaluation harnesses, verifiers, benchmark development, or reproducible experimentation.
  • Knowledge of GPU infrastructure, distributed training, Kubernetes, workload schedulers, cluster capacity management, performance analysis, open-source contributor workflows, release readiness, or operational metrics.

You will also be eligible for equity and benefits.