# Engineering Manager, DGX Cloud Production Engineering

**Company**: NVIDIA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Remote/Engineering-Manager--DGX-Cloud-Production-Engineering_JR2021537?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_589ed95e-092

## Description

NVIDIA's DGX Cloud Production Engineering team is seeking an Engineering Manager to lead a dynamic group of engineers. This role involves leading a team of software and production engineers in building and operating the DGX Cloud infrastructure across NVIDIA Cloud Partner (NCP) and on-prem environments.

Responsibilities:

- Lead a team of software and production engineers in building and operating the DGX Cloud infrastructure.

- Drive execution in cluster operations, Kubernetes operability, automation, GitOps, observability, and incident response.

- Define team priorities, roadmap, staffing, and operational ownership.

- Partner with platform, workload, storage, networking, security, and TPM teams to improve production readiness.

- Encourage a healthy on-call and incident review culture focused on learning, ownership, and durable fixes.

- Mentor engineers, grow technical leaders, and craft clear ownership across ambiguous problem spaces.

Requirements:

- 8+ overall years of industry experience, including 2+ years leading or managing engineers.

- Proven experience in building or operating production infrastructure, cloud platforms, Kubernetes environments, or distributed systems.

- Strong understanding of reliability engineering, automation, observability, incident response, and operational excellence.

- Ability to work effectively across teams and influence without direct authority.

- Clear communication, strong prioritization, and good judgment in fast-paced environments.

- BS/MS in Computer Science or equivalent experience.

Preferred qualifications:

- Experience leading SRE, production engineering, infrastructure automation, or platform teams.

- Familiarity with GPU infrastructure, Kubernetes fleet operations, GitOps, BMaaS/VMaaS, managed Kubernetes, or multi-cloud environments.

- A track record of reducing toil, improving SLOs, and turning operational work into automated systems.

NVIDIA offers highly competitive salaries and a comprehensive benefits package.

## Skills

### Required
- Kubernetes
- cloud platforms
- distributed systems
- reliability engineering
- automation
- observability
- incident response

### Nice to have
- GPU infrastructure
- Kubernetes fleet operations
- GitOps
- BMaaS/VMaaS
- managed Kubernetes
- multi-cloud environments

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Remote/Engineering-Manager--DGX-Cloud-Production-Engineering_JR2021537?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
