Description
NVIDIA is looking for a Distinguished Engineer to act as a senior technical leader in the Production Engineering group, focusing on cluster operations involving DGX Cloud GPU capacity.
At NVIDIA, Production Engineering ensures large-scale production systems are reliable, straightforward to lead, and increasingly automated across NVIDIA's DGX Cloud resources. The role centres on the operational framework supporting DGX Cloud environments spanning on-premises, major cloud providers, and NVIDIA Cloud Partner locations.
Responsibilities include:
- Defining the long-range technical strategy for operating DGX Cloud clusters consistently across on-prem, hyperscalers, and NeoCloud environments
- Defining the architectural vision and core operational guidelines for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability throughout DGX Cloud resources
- Guiding the roadmap and execution of critical cross-organizational investments that improve production readiness, operational safety, performance, and cross-team coordination
- Making and guiding high-impact technical decisions that resolve how platform, hardware, provider, and service teams coordinate to operate DGX Cloud resources in production
- Developing robust workflows, interfaces, and engineering collaboration across Kubernetes production service, provider and hardware readiness, on-prem and bare-metal infrastructure operations, and service-layer reliability domains
Requirements:
- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience
- 18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments
- Confirmed company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms
- Confirmed experience in establishing operating models, architectural direction, and engineering standards across various technical domains and organizations
- Consistent record leading large, cross-team technical efforts from concept through production, including aligning collaborators, navigating complexity and delivering measurable outcomes
Benefits:
- Equity
- Other benefits
Applications for this job will be accepted at least until September 6, 2026.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Distinguished-Engineer--Production-Engineering--Data-Center-Automation_JR2024978