# Engineering Manager, LLM Performance

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Engineering-Manager--LLM-Performance_JR2019950?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_2dbda25c-40e

## Description

At NVIDIA, we're accelerating the AI revolution, particularly in large language models (LLMs) and vision language models (VLMs). We're seeking an Engineering Manager to lead the development of next-generation LLM/VLM/VLA inference software technologies.

**Job Summary:** This is a high-impact, hands-on leadership role that requires deep technical expertise and world-class management skills. You will lead and grow a team of engineers who are pushing the performance of LLM inference across multiple frameworks, including TensorRT LLM, vLLM, SGLang, and Dynamoo on datacenter products.

**Responsibilities:**

- Lead and grow a team responsible for pushing the performance of LLM inference across multiple LLM frameworks.

- Drive the design, implementation, and optimization of features key to performance in LLM inference.

- Continuously improve the performance of LLM inference on current and upcoming NVIDIA datacenter architectures and GPUs.

- Improve the performance of LLM inference of important foundation models.

- Work with inference benchmark teams to help tune performance for key workloads.

- Integrate cutting-edge technologies developed at NVIDIA and offer an intuitive developer experience for LLM deployment.

- Lead software development execution, with responsibility for project planning, milestone delivery, and cross-functional coordination.

**Requirements:**

- MS, PhD, or equivalent experience in Computer Science, Computer Engineering, AI, or a related technical field.

- 7+ overall years of software engineering experience, including 3+ years of technical leadership experience.

- Proven ability to lead and scale high-performing engineering teams.

- Strong background in C++ or Python, with expertise in software design and delivering production-quality software libraries.

- Demonstrated expertise in large language models (LLM) and/or vision language models (VLM) and/or inference in general.

**Benefits:** You will also be eligible for equity and benefits.

**How to Apply:** Applications for this job will be accepted at least until June 27, 2026.

## Skills

### Required
- C++
- Python
- Computer Science
- AI
- LLM
- VLM
- inference

### Nice to have
- GPU architecture
- CUDA programming
- system-level performance tuning
- TensorRT-LLM
- vLLM
- SGLang

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Engineering-Manager--LLM-Performance_JR2019950?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
