Description
We are seeking a Principal Engineer to join our CSP Engagements team as the technical focal point for end-to-end performance. You will work directly with engineering teams of key CSP/hyperscale customers to ensure they achieve various performance targets on NVIDIA platforms.
In this role, you will:
- Drive performance characterization work streams with engineering teams of key CSP/hyperscale customers , ensuring they understand platform performance expectations, profiling methodology, and tuning options for their specific workloads
- Gather and synthesize CSP performance feedback , identify gaps between expected and actual throughput, and champion optimization priorities back into NVIDIA's CUDA, NCCL, driver, and firmware teams
- Ensure key open-source performance and stress tools are updated and validated for the latest NVIDIA rack-scale systems, GPU architectures, and CPU platforms
- Work closely with CSPs to ensure their own performance and validation tooling reflects the latest GPU capabilities, memory hierarchy changes, and platform-specific tuning parameters
- Conduct cross-CSP performance comparison and pattern analysis , identify configuration, software, or workload differences that explain performance gaps between deployments
- Collaborate with CSPs to ensure performance-related integration work is ready ahead of deployment milestones
- Define test strategies and tooling requirements for performance validation , both for NVIDIA internal certification and customer acceptance
Requirements:
- 15+ years of experience in systems performance engineering, ideally in GPU/HPC/ML infrastructure
- Proficiency in GPU workload profiling: nsight systems, nsight compute, DCGM metrics, or equivalent instrumentation
- Understanding of distributed training performance dynamics: computation/communication overlap, pipeline bubbles, memory bandwidth utilization, collective efficiency
- Statistical methods for performance analysis: regression detection, confidence intervals, A/B comparison at scale
- Understanding of how the full software stack impacts performance: driver overhead, collective algorithm selection, memory allocation, scheduling, firmware power management
- Strong data analysis and visualization skills (Python, pandas, dashboards)
- Customer obsession , genuine passion for understanding why customers aren't achieving expected performance and driving solutions
- Ability to communicate performance findings to both deep technical audiences and executive leadership
- Demonstrated success influencing multiple engineering teams to prioritize performance improvements
Preferred qualifications:
- Experience profiling and optimizing distributed training at 1000+ GPU scale
- Background in ML infrastructure performance at a CSP/hyperscaler
- Familiarity with NVIDIA platforms and profiling tools
- Experience building automated performance regression detection systems for production environments
- Understanding of inference workload performance dynamics
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer--E2E-Performance-and-Goodput---CSP-Engagements_JR2020321