New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Site Reliability Engineering (SRE)

NVIDIA
Apply →
hybrid senior full-time 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333

First indexed 9 Sept 2026

Description

Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline that designs, builds, and maintains large-scale production systems with high efficiency and availability.

You will lead the technical strategy and roadmap for large-scale, cross-functional SRE initiatives that improve reliability, scalability, and developer productivity across enterprise systems.

Responsibilities:

  • Design, build, and maintain resilient distributed systems that power NVIDIA's next-generation AI-driven enterprise products and services.
  • Architect and develop AI Agents, AI Skills to accelerate platform operations.
  • Drive automation and observability improvements, using metrics and analytics to enhance performance, reliability, and efficiency.
  • Collaborate across Cloud, Platform, Security, and AI/ML teams to implement modern SRE components that ensure high availability and secure operations.
  • Analyze and troubleshoot complex systems, championing best practices in system design, incident management, and postmortem analysis.
  • Mentor and influence engineers across teams, fostering technical excellence and a culture of reliability engineering.

Requirements:

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
  • BS degree in Computer Science or a related technical field involving coding.
  • Strong proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code.
  • Experience with infrastructure-as-code such as AWS CDK, AWS CloudFormation, Terraform, or CrossPlane.
  • Solid understanding of OpenTelemetry or other Observability implementation at scale.
  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP).
  • Outstanding problem-solving, communication, and teamwork skills, with the ability to influence across technical and interpersonal boundaries.

Benefits:

  • Base salary range: 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.
  • Equity and benefits package.