New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Distinguished Engineer, Scaled Out Inferencing

NVIDIA
Apply →
remote senior full-time Santa Clara, CA

First indexed 11 Aug 2026

Description

As a technology leader at NVIDIA, you will lead the development of our global strategy for scaled-out AI inferencing. You will architect high-throughput, low-latency distributed pipelines and model serving strategies required for massive scale and production reliability.

Responsibilities:

  • Architect distributed pipelines, define and drive the technical implementation of high-throughput, low-latency, distributed inference systems to support massive-scale AI workloads.
  • Collaborate on hardware-software co-optimization, drive performance tuning at the kernel and driver level, optimizing GPU resource management and hardware acceleration for production-grade model serving.
  • Guide and influence open source projects Dynamo, TensorRT-LLM, and ecosystem projects (vLLM, SGLang, Linux, Kubernetes, Ray) to bring state-of-the-art inferencing on NVIDIA accelerated hardware.
  • Orchestrate model lifecycles, lead the strategy for full-lifecycle model management, including automated deployment, versioning, and intelligent scaling across varied cloud and datacenter environments.
  • Collaborate with customers, infrastructure providers, and partners to ensure NVIDIA’s solutions set the industry standard for performance and availability.
  • Lead all technical aspects of planning and continuous evolution of a large technical scope.

Requirements:

  • 16+ overall years in technical roles with a recent long-term focus on AI infrastructure and more recent direct experience in large-scale inference orchestration.
  • 7-10+ years of leadership experience.
  • BS/MS or higher or equivalent experience in systems/software engineering, or related engineering fields.
  • Deep technical expertise: proficiency in GPU architecture, hardware acceleration, and low-level performance tuning (CUDA, kernels) alongside cloud-native architectures for multi-tenant model serving.
  • Proven success delivering high-impact technically complex solutions that achieve high levels of transparency into resource utilization, performance, and operational insights.
  • Technical leadership: develop and advance consensus and organizational alignment across technical leadership and the highest level of senior corporate leadership.
  • Strong collaboration and influence skills, capable of leading engineering engagement, communicating with peers, partners, and working with high-performance and accelerated computing customers.

Benefits:

  • Competitive salaries
  • Generous benefits package (visit www.nvidiabenefits.com)
  • Equity eligibility