Description
As a technology leader at NVIDIA, you will lead the development of our global strategy for scaled-out AI inferencing. You will architect high-throughput, low-latency distributed pipelines and model serving strategies required for massive scale and production reliability.
Responsibilities:
- Architect distributed pipelines, define and drive the technical implementation of high-throughput, low-latency, distributed inference systems to support massive-scale AI workloads.
- Collaborate on hardware-software co-optimization, drive performance tuning at the kernel and driver level, optimizing GPU resource management and hardware acceleration for production-grade model serving.
- Guide and influence open source projects Dynamo, TensorRT-LLM, and ecosystem projects (vLLM, SGLang, Linux, Kubernetes, Ray) to bring state-of-the-art inferencing on NVIDIA accelerated hardware.
- Orchestrate model lifecycles, lead the strategy for full-lifecycle model management, including automated deployment, versioning, and intelligent scaling across varied cloud and datacenter environments.
- Collaborate with customers, infrastructure providers, and partners to ensure NVIDIA’s solutions set the industry standard for performance and availability.
- Lead all technical aspects of planning and continuous evolution of a large technical scope.
Requirements:
- 16+ overall years in technical roles with a recent long-term focus on AI infrastructure and more recent direct experience in large-scale inference orchestration.
- 7-10+ years of leadership experience.
- BS/MS or higher or equivalent experience in systems/software engineering, or related engineering fields.
- Deep technical expertise: proficiency in GPU architecture, hardware acceleration, and low-level performance tuning (CUDA, kernels) alongside cloud-native architectures for multi-tenant model serving.
- Proven success delivering high-impact technically complex solutions that achieve high levels of transparency into resource utilization, performance, and operational insights.
- Technical leadership: develop and advance consensus and organizational alignment across technical leadership and the highest level of senior corporate leadership.
- Strong collaboration and influence skills, capable of leading engineering engagement, communicating with peers, partners, and working with high-performance and accelerated computing customers.
Benefits:
- Competitive salaries
- Generous benefits package (visit www.nvidiabenefits.com)
- Equity eligibility
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Distinguished-Engineer--Scaled-Out-Inferencing_JR2023118