# Site Reliability Engineering (SRE)

**Company**: NVIDIA
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Salary**: 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Staff-Site-Reliability-Engineer---AI-Platform-Runtime_JR2023969?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_50c9f536-024

## Description

Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline that designs, builds, and maintains large-scale production systems with high efficiency and availability.

You will lead the technical strategy and roadmap for large-scale, cross-functional SRE initiatives that improve reliability, scalability, and developer productivity across enterprise systems.

Responsibilities:

- Design, build, and maintain resilient distributed systems that power NVIDIA's next-generation AI-driven enterprise products and services.

- Architect and develop AI Agents, AI Skills to accelerate platform operations.

- Drive automation and observability improvements, using metrics and analytics to enhance performance, reliability, and efficiency.

- Collaborate across Cloud, Platform, Security, and AI/ML teams to implement modern SRE components that ensure high availability and secure operations.

- Analyze and troubleshoot complex systems, championing best practices in system design, incident management, and postmortem analysis.

- Mentor and influence engineers across teams, fostering technical excellence and a culture of reliability engineering.

Requirements:

- 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.

- BS degree in Computer Science or a related technical field involving coding.

- Strong proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code.

- Experience with infrastructure-as-code such as AWS CDK, AWS CloudFormation, Terraform, or CrossPlane.

- Solid understanding of OpenTelemetry or other Observability implementation at scale.

- Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP).

- Outstanding problem-solving, communication, and teamwork skills, with the ability to influence across technical and interpersonal boundaries.

Benefits:

- Base salary range: 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.

- Equity and benefits package.

## Skills

### Required
- Python
- Typescript
- JavaScript
- Go
- AWS CDK
- AWS CloudFormation
- Terraform
- CrossPlane
- OpenTelemetry
- Kubernetes
- public cloud services

### Nice to have
- Public Cloud
- large-scale automation systems

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Staff-Site-Reliability-Engineer---AI-Platform-Runtime_JR2023969?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
