# Senior Software Engineer, Resilience Engineering - DGX Cloud

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer--Resilience-Engineering---DGX-Cloud_JR2021860?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_ea96e7a2-967

## Description

NVIDIA is seeking a Senior Software Engineer - Resilience Engineering to join the DGX Cloud team. As a Senior Software Engineer, you will be a pivotal part of a team that redefines operational excellence.

**Job Description:**

You will build the organisation-wide reliability strategy, guiding how NVIDIA matures its operational practices in a 24/7 environment. You will stand up a rigorous SLO program, defining and maintaining high standards across teams. You will lead incident response for high-severity incidents, ensuring low drama and high signal resolution.

Your responsibilities will include:

- Building and improving production code daily, enhancing the data platform and related tooling.

- Implementing chaos engineering, failure injection, and resilience testing to elevate the team's standard practices.

- Improving standards by setting an example with your hands-on experience and leadership.

**Requirements:**

- Deep, hands-on experience running large-scale production systems with a proven track record.

- A detailed understanding of failure modes in large systems, including cascading dependencies and retry storms.

- Strong software engineering skills with current, hands-on experience in Go, Python, or similar languages.

- Proven experience in establishing and maintaining an SLO program with operational rigor.

- Practical experience in reliability fields such as chaos engineering and failure injection.

- The ability to influence across team boundaries through credibility and expertise.

- 8+ years of industry experience.

- Bachelor's or Master's degree, or equivalent experience operating systems at scale.

**Nice to Have:**

- Experience within a world-class reliability function like Google SRE or Meta production engineering.

- Expertise in operating GPU, HPC, or AI training infrastructure with outstanding failure modes.

- A track record of measurable reliability improvements within an organisation.

- Proficiency with modern observability and operational tools like Prometheus, OpenTelemetry, Grafana, PagerDuty, and Rootly.

NVIDIA offers highly competitive salaries and a comprehensive benefits package.

## Skills

### Required
- Go
- Python
- chaos engineering
- failure injection
- resilience testing
- SLO program
- incident response

### Nice to have
- Google SRE
- Meta production engineering
- GPU
- HPC
- AI training infrastructure
- Prometheus
- OpenTelemetry
- Grafana
- PagerDuty
- Rootly

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer--Resilience-Engineering---DGX-Cloud_JR2021860?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
