# Senior Reliability Engineer, DGX Cloud

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Reliability-Engineer--DGX-Cloud_JR2019933-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_10ca90f7-309

## Description

Join NVIDIA as a Senior Reliability Engineer, DGX Cloud, and be pivotal in redefining operational excellence. You will build an organisation-wide reliability strategy, guide NVIDIA's operational practices in a 24/7 environment, and lead incident response for high-severity incidents.

**Responsibilities:**

- Develop and implement a rigorous Service Level Objective (SLO) program with high standards across teams.

- Lead incident response for high-severity incidents, ensuring efficient resolution.

- Enhance production code daily, improving the data platform and related tooling.

- Implement chaos engineering, failure injection, and resilience testing.

- Set an example with hands-on experience and leadership, improving standards.

**Requirements:**

- Hands-on experience running large-scale production systems with a proven track record.

- Understanding of failure modes in large systems, including cascading dependencies and retry storms.

- Strong software engineering skills in languages like Go, Python.

- Experience establishing and maintaining an SLO program.

- Practical experience in reliability fields such as chaos engineering and failure injection.

- Ability to influence across team boundaries through credibility and expertise.

- 10+ years of industry experience with a Bachelor's or Master's degree, or equivalent experience.

**Benefits:**

- Competitive salaries

- Comprehensive benefits package

- Equity eligibility

- Visit www.nvidiabenefits.com/ for more information.

## Skills

### Required
- large-scale systems
- reliability engineering
- SLO program
- incident response
- chaos engineering
- Go
- Python

### Nice to have
- Google SRE
- Meta production engineering
- GPU
- HPC
- AI training infrastructure
- Prometheus
- OpenTelemetry
- Grafana
- PagerDuty
- Rootly

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Reliability-Engineer--DGX-Cloud_JR2019933-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
