Description
NVIDIA's Local AI team is seeking a System Software Manager to lead development of an efficient on-device AI software stack. The software stack will support RTX, RTX Pro, and DGX-class systems. This role focuses on high-performance local inference, agentic workloads, low latency, efficient memory use, scalable infrastructure, practical deployment on resource-constrained platforms, and delivering a streamlined out-of-box experience for developers and end users.
Responsibilities:
- Lead and grow a team building the on-device AI inference platform for RTX, RTX Pro, and DGX GPUs, with accountability for execution, technical direction, delivery quality, and roadmap alignment.
- Drive cross-functional alignment with NVIDIA’s software, research, architecture, and product teams, along with industry partners and open-source communities, to build strategy and strengthen the AI ecosystem across RTX and DGX platforms.
- Provide technical leadership for the architecture and evolution of modern inference runtimes and execution stacks across frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX, spanning workloads including LLMs, vision-language models, TTS, ASR, and diffusion models.
- Mentor engineers, develop technical leaders, and foster a high-performance team culture centred on innovation, collaboration, and operational excellence.
- Coordinate end-to-end optimization of AI models, data pipelines, and inference runtimes to improve performance across current and next-generation GPU architectures.
- Drive adoption of model optimization techniques such as quantization, pruning, sparsity, and distillation to enable efficient deployment of large models on local and edge devices.
- Establish team processes for system-level debugging, performance optimization, and performance-accuracy trade-off analysis, including infrastructure for performance and accuracy sweeps, gap analysis, and production-readiness improvements.
Requirements:
- 5+ overall years of industry experience and 2+ years of engineering leadership experience, combined with a Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or a related field.
- Proven experience leading high-performing engineering teams in systems software, AI infrastructure, inference runtimes, or related domains.
- Strong technical foundation in C++ software development, debugging, data structures, algorithms, and machine learning systems.
- Extensive background in AI inference pipelines and Deep Learning frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
- Deep understanding of inference backends and runtime internals, including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
- Strong analytical and problem-solving skills, with the ability to balance technical depth, execution speed, and organizational priorities in a fast-paced environment.
- Excellent written and verbal communication skills, with proven ability to collaborate across engineering, product, research, and executive collaborators.
Benefits:
- Competitive salaries
- Generous benefits package
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Pune/Manager--System-Software-Engineering---Local-AI_JR2021898-1