New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Engineering Manager, LLM Performance

NVIDIA
Apply →
hybrid senior full-time Santa Clara, CA

First indexed 24 Jun 2026

Description

At NVIDIA, we're accelerating the AI revolution, particularly in large language models (LLMs) and vision language models (VLMs). We're seeking an Engineering Manager to lead the development of next-generation LLM/VLM/VLA inference software technologies.

Job Summary: This is a high-impact, hands-on leadership role that requires deep technical expertise and world-class management skills. You will lead and grow a team of engineers who are pushing the performance of LLM inference across multiple frameworks, including TensorRT LLM, vLLM, SGLang, and Dynamoo on datacenter products.

Responsibilities:

  • Lead and grow a team responsible for pushing the performance of LLM inference across multiple LLM frameworks.
  • Drive the design, implementation, and optimization of features key to performance in LLM inference.
  • Continuously improve the performance of LLM inference on current and upcoming NVIDIA datacenter architectures and GPUs.
  • Improve the performance of LLM inference of important foundation models.
  • Work with inference benchmark teams to help tune performance for key workloads.
  • Integrate cutting-edge technologies developed at NVIDIA and offer an intuitive developer experience for LLM deployment.
  • Lead software development execution, with responsibility for project planning, milestone delivery, and cross-functional coordination.

Requirements:

  • MS, PhD, or equivalent experience in Computer Science, Computer Engineering, AI, or a related technical field.
  • 7+ overall years of software engineering experience, including 3+ years of technical leadership experience.
  • Proven ability to lead and scale high-performing engineering teams.
  • Strong background in C++ or Python, with expertise in software design and delivering production-quality software libraries.
  • Demonstrated expertise in large language models (LLM) and/or vision language models (VLM) and/or inference in general.

Benefits: You will also be eligible for equity and benefits.

How to Apply: Applications for this job will be accepted at least until June 27, 2026.

This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Engineering-Manager--LLM-Performance_JR2019950