# Principal Software Engineer, DGX Cloud Production Engineering

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer--DGX-Cloud-Production-Engineering_JR2022554?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_9518f659-43f

## Description

NVIDIA is seeking a principal-level software engineer to build the next generation of our Kubernetes platform. Our teams build foundational capabilities for self-service GPU infrastructure, managed Kubernetes control planes, management of cluster operations, and automation for large-scale AI and improved computational environments.

In this role, you will work at the intersection of platform architecture and production-grade Kubernetes lifecycle systems. You will lead technical strategy and execution for systems that make clusters easier to provision, upgrade, operate, and scale across cloud and on-premises environments.

**Responsibilities:**

- Lead the architecture and development of core Kubernetes platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations.

- Design and build highly reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.

- Define technical requirements, validation criteria, production-readiness practices, and the direction for declarative workflows and automation across the Kubernetes stack.

- Collaborate across engineering teams to create cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency.

- Lead the diagnosis and resolution of complex platform issues spanning infrastructure, runtime, networking, hardware, and operations, improving the scalability, resilience, and operability of systems supporting large-scale AI deployments.

- Influence engineering standards, architectural decisions, and long-term platform strategy.

- Mentor senior engineers and raise the bar for design quality, execution, and engineering rigor across the organization.

**Requirements:**

- BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.

- 15+ years of relevant software engineering experience, including experience building and operating large-scale production systems.

- Deep expertise in Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.

- A strong background in distributed systems design, reliability, scalability, and failure recovery.

- Proven experience building platform software, infrastructure control planes, or foundations for managed services.

- Strong programming skills in one or more systems or cloud-native languages, such as Go, Python, Rust, or C++.

- Experience designing clear APIs and abstractions for platform consumers and engineering teams.

- Demonstrated ability to provide technical leadership across team boundaries and drive ambiguous, cross-functional initiatives to completion.

- Excellent communication and collaboration skills, backed by a sustained record of significant technical contributions and recognized expertise influencing department-level architecture and high-priority company initiatives.

**Benefits:**

- Competitive salaries

- Generous benefits package

- Equity

Applications for this job will be accepted at least until August 15, 2026.

## Skills

### Required
- Kubernetes
- distributed systems design
- cloud-native languages
- platform software
- infrastructure control planes

### Nice to have
- fleet management
- cluster upgrades
- node lifecycle
- remediation
- day-2 operations
- declarative infrastructure
- Kubernetes controllers
- GitOps
- policy-driven platform automation
- public-cloud and bare-metal infrastructure environments
- AI
- GPU
- HPC
- large-scale accelerated computing platforms

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer--DGX-Cloud-Production-Engineering_JR2022554?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
