New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Software Engineer, Resilience Engineering - DGX Cloud

NVIDIA
Apply →
remote senior full-time Santa Clara, CA

First indexed 29 Jul 2026

Description

NVIDIA is seeking a Senior Software Engineer - Resilience Engineering to join the DGX Cloud team. As a Senior Software Engineer, you will be a pivotal part of a team that redefines operational excellence.

Job Description:

You will build the organisation-wide reliability strategy, guiding how NVIDIA matures its operational practices in a 24/7 environment. You will stand up a rigorous SLO program, defining and maintaining high standards across teams. You will lead incident response for high-severity incidents, ensuring low drama and high signal resolution.

Your responsibilities will include:

  • Building and improving production code daily, enhancing the data platform and related tooling.
  • Implementing chaos engineering, failure injection, and resilience testing to elevate the team's standard practices.
  • Improving standards by setting an example with your hands-on experience and leadership.

Requirements:

  • Deep, hands-on experience running large-scale production systems with a proven track record.
  • A detailed understanding of failure modes in large systems, including cascading dependencies and retry storms.
  • Strong software engineering skills with current, hands-on experience in Go, Python, or similar languages.
  • Proven experience in establishing and maintaining an SLO program with operational rigor.
  • Practical experience in reliability fields such as chaos engineering and failure injection.
  • The ability to influence across team boundaries through credibility and expertise.
  • 8+ years of industry experience.
  • Bachelor's or Master's degree, or equivalent experience operating systems at scale.

Nice to Have:

  • Experience within a world-class reliability function like Google SRE or Meta production engineering.
  • Expertise in operating GPU, HPC, or AI training infrastructure with outstanding failure modes.
  • A track record of measurable reliability improvements within an organisation.
  • Proficiency with modern observability and operational tools like Prometheus, OpenTelemetry, Grafana, PagerDuty, and Rootly.

NVIDIA offers highly competitive salaries and a comprehensive benefits package.