New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Reliability Engineer, DGX Cloud

NVIDIA
Apply →
remote senior full-time Santa Clara, CA

First indexed 23 Jun 2026

Description

Join NVIDIA as a Senior Reliability Engineer, DGX Cloud, and be pivotal in redefining operational excellence. You will build an organisation-wide reliability strategy, guide NVIDIA's operational practices in a 24/7 environment, and lead incident response for high-severity incidents.

Responsibilities:

  • Develop and implement a rigorous Service Level Objective (SLO) program with high standards across teams.
  • Lead incident response for high-severity incidents, ensuring efficient resolution.
  • Enhance production code daily, improving the data platform and related tooling.
  • Implement chaos engineering, failure injection, and resilience testing.
  • Set an example with hands-on experience and leadership, improving standards.

Requirements:

  • Hands-on experience running large-scale production systems with a proven track record.
  • Understanding of failure modes in large systems, including cascading dependencies and retry storms.
  • Strong software engineering skills in languages like Go, Python.
  • Experience establishing and maintaining an SLO program.
  • Practical experience in reliability fields such as chaos engineering and failure injection.
  • Ability to influence across team boundaries through credibility and expertise.
  • 10+ years of industry experience with a Bachelor's or Master's degree, or equivalent experience.

Benefits:

  • Competitive salaries
  • Comprehensive benefits package
  • Equity eligibility
  • Visit www.nvidiabenefits.com/ for more information.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Reliability-Engineer--DGX-Cloud_JR2019933-1