# Senior Product Manager – AI Inference Performance

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Product-Manager---AI-Inference-Performance_JR2022876?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_a3ae1859-57e

## Description

NVIDIA is seeking a highly technical Product Manager to own products that help customers extract the best possible performance from AI models and applications running on NVIDIA hardware. The role involves turning deep optimization techniques into products that a broad range of customers can adopt, spanning the entire inference stack.

**Key Responsibilities:**

- Own the inference performance roadmap, setting direction across the stack.

- Build platforms that generalize across model families, deployment topologies, and customer sizes.

- Define the performance strategy for agentic applications and drive capabilities around cross-turn cache reuse and efficient handling of idle time.

- Develop the framework and ecosystem strategy, partnering with open-source communities and internal engineering teams.

- Own benchmarking and performance claims, defining methodologies and metrics.

**Requirements:**

- 12+ years in product management at a technology company, or comparable experience.

- Depth in AI inference optimization, including KV caching, quantization, and speculative decoding.

- Familiarity with inference and orchestration frameworks such as TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.

- Proven track record of working independently and driving strategies to shipped outcomes.

- Operational experience running a live product, including release management and customer support.

**Preferred Qualifications:**

- Engineering experience with LLM inference performance optimization.

- Open-source contributions or product leadership in relevant projects.

- Habit of reading relevant research and translating it into roadmap decisions.

NVIDIA offers highly competitive salaries and a comprehensive benefits package. Applications will be accepted until August 17, 2026.

## Skills

### Required
- AI inference optimization
- Product management
- TensorRT-LLM
- vLLM
- SGLang
- NVIDIA Dynamo

### Nice to have
- LLM inference performance optimization
- Open-source contributions
- Research translation

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Product-Manager---AI-Inference-Performance_JR2022876?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
