# Engineering Manager, Deep Learning Inference

**Company**: NVIDIA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Engineering-Manager--Deep-Learning-Inference_JR2022350?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_37be1038-e9e

## Description

NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment.

You will shape the software powering today’s most sophisticated AI systems , from large language models to multimodal generative AI , all accelerated on NVIDIA GPUs.

The Deep Learning Inference team develops and optimizes open-source frameworks that make AI deployment scalable, efficient, and accessible , including SGLang, vLLM, and FlashInfer.

Our work enables developers worldwide to harness NVIDIA accelerators for real-time inference at every scale, from datacenter clusters to edge devices.

### Responsibilities

- Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.

- Guide the strategy, roadmap, and execution of NVIDIA's OSS inference frameworks engineering.

- Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.

- Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.

- Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).

- Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA’s broader AI and software strategies.

- Foster a culture of technical excellence, open collaboration, and continuous innovation.

### Requirements

- MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.

- 6+ overall years of software development experience, including 3+ years in technical leadership or engineering management.

- Strong background in C/C++ software design and development; proficiency in Python is a plus.

- Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization.

- Proven record of deploying or optimizing deep learning models in production environments.

- Experience leading teams using Agile or collaborative software development practices.

### Ways to Stand out from The Crowd

- Significant open-source contributions to deep learning or inference frameworks such as PyTorch, vLLM, SGLang, Triton, or TensorRT-LLM.

- Deep understanding of multi-GPU communications (NIXL, NCCL, NVSHMEM) and distributed inference architectures.

- Expertise in performance modeling, profiling, and system-level optimization across CPU and GPU platforms.

- Proven ability to mentor engineers, guide architectural decisions, and deliver complex projects with measurable impact.

- Publications, patents, or talks on LLM serving, model optimization, or GPU performance engineering.

### Benefits

- Highly competitive salaries

- Comprehensive benefits package

- Equity eligibility

## Skills

### Required
- C/C++
- Python
- GPU programming
- CUDA
- Triton
- CUTLASS
- performance optimization
- deep learning models
- Agile software development

### Nice to have
- open-source contributions
- multi-GPU communications
- distributed inference architectures
- performance modeling
- profiling
- system-level optimization
- mentor engineers

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Engineering-Manager--Deep-Learning-Inference_JR2022350?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
