Description
Join NVIDIA as a Senior Reliability Engineer, DGX Cloud, and be pivotal in redefining operational excellence. You will build an organisation-wide reliability strategy, guide NVIDIA's operational practices in a 24/7 environment, and lead incident response for high-severity incidents.
Responsibilities:
- Develop and implement a rigorous Service Level Objective (SLO) program with high standards across teams.
- Lead incident response for high-severity incidents, ensuring efficient resolution.
- Enhance production code daily, improving the data platform and related tooling.
- Implement chaos engineering, failure injection, and resilience testing.
- Set an example with hands-on experience and leadership, improving standards.
Requirements:
- Hands-on experience running large-scale production systems with a proven track record.
- Understanding of failure modes in large systems, including cascading dependencies and retry storms.
- Strong software engineering skills in languages like Go, Python.
- Experience establishing and maintaining an SLO program.
- Practical experience in reliability fields such as chaos engineering and failure injection.
- Ability to influence across team boundaries through credibility and expertise.
- 10+ years of industry experience with a Bachelor's or Master's degree, or equivalent experience.
Benefits:
- Competitive salaries
- Comprehensive benefits package
- Equity eligibility
- Visit www.nvidiabenefits.com/ for more information.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Reliability-Engineer--DGX-Cloud_JR2019933-1