Description
NVIDIA's DGX Cloud organization seeks a Senior Systems Software Engineer to tackle cutting-edge hardware and software innovation for accelerated computing in AI workloads. You will focus on scaling AI infrastructure, minimizing total cost of ownership, and enabling future AI innovation.
Job Overview
As a Senior Systems Software Engineer, you will lead end-to-end performance and scalability analysis across the Kubernetes-based accelerated runtime stack. You will design and contribute to upstream architectural changes to enable reliable operation at hyperscale cluster sizes. Your goal will be to improve container startup and cold-start latency, assess and improve open-source projects, and advance scalability and performance of confidential containers on Kubernetes.
Responsibilities
- Lead end-to-end performance and scalability analysis across the Kubernetes-based accelerated runtime stack
- Design and contribute upstream architectural changes to the Kubernetes control plane and related projects
- Improve container startup and cold-start latency to enable smooth, low-latency inference scaling on Kubernetes
- Assess, improve, and contribute to open-source projects that make Kubernetes an outstanding platform for AI workloads
- Advance scalability and performance of confidential containers on Kubernetes
- Use DSX and related large-scale simulation infrastructure to model full AI-factory deployments and validate scalability
- Collaborate with AI researchers, developers, customers, and upstream communities to design automated, at-scale workload tests
- Document methods and results clearly and present findings internally and at industry events
Requirements
- Bachelor's or Master's degree in Engineering or equivalent experience
- 8+ years of experience in computer architecture, networking, storage systems, and accelerator-based platforms
- Expertise in Kubernetes and familiarity with the broader CNCF ecosystem
- Deep experience with large-scale, parallel, distributed accelerator systems and performance optimization of AI workloads
- Experience with performance modeling and benchmarking for large-scale systems
- Proficiency in Golang and/or Python
- Strong familiarity with the NVIDIA software stack across training and inference
- Expertise with at least one major public cloud provider
Benefits
You will also be eligible for equity and benefits.