# Senior Staff Site Reliability Engineer

**Company**: NVIDIA
**Location**: Bengaluru
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2024618?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_b7dedb1e-c35

## Description

NVIDIA is looking for a Senior Staff Site Reliability Engineer to join our team in India. This is a senior individual-contributor role that combines real-time incident leadership with deep hands-on engineering, focused on how we detect, respond to, and prevent issues at scale across NVIDIA's AI-powered enterprise platforms.

You will operate as an Incident Commander during critical events, set the technical direction for reliability engineering across multiple teams, and build the automation, observability, and AI-assisted tooling that reduce operational toil and raise reliability over time. You will also mentor and grow SRE talent as we scale the practice in India.

**Responsibilities:**

- Lead major incidents end to end , driving triage, cross-team coordination, decision making, and executive communication across global time zones.

- Set technical direction for SRE initiatives that improve reliability, scalability, and developer efficiency across NVIDIA's enterprise systems, and drive them to adoption beyond your immediate team.

- Design, build, and operate distributed systems , including Kubernetes-based and cloud-native infrastructure , that power NVIDIA's AI-powered enterprise products and services.

- Build automation for incident detection, triage, communication, and remediation, replacing manual runbooks with self-healing systems.

- Improve observability and signal quality to enable earlier detection, reduce alert noise, and eliminate reliance on user-reported issues.

- Drive root cause analysis and translate learnings into systemic fixes, automation, and prevention mechanisms; raise the bar on post-incident review quality across the org.

- Apply AI and data-driven techniques , LLMs, anomaly detection, signal correlation , to enhance incident triage, summarisation, and decision support.

- Champion AI-assisted engineering practices, including coding agents and LLM-powered tooling, to accelerate day-to-day engineering workflows.

- Partner with Cloud, Platform, Security, and AI/ML teams to embed SRE best practices, define SLOs and error budgets, and influence architecture early in the design cycle.

- Mentor engineers, raise engineering standards through design and code review, and help build a strong reliability culture in the India organisation.

**Requirements:**

- 10+ years of experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management roles, with a track record of technical leadership at scale.

- BS or MS degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).

- Proven experience acting as an Incident Commander or leading major incident response in complex, high-availability environments.

- Deep understanding of distributed systems, monitoring, and reliability engineering principles , SLIs/SLOs, error budgets, capacity planning, and graceful degradation.

- Strong proficiency in at least one programming language (e.g., Python, Go, Java ) to build production-grade automation and tooling.

- Hands-on expertise with public cloud platforms (AWS, Azure, or GCP) and container technologies such as Docker and Kubernetes.

- Solid experience with infrastructure-as-code tooling (e.g., Terraform, AWS CDK, CloudFormation) and CI/CD pipelines.

- Strong Linux/Unix and networking fundamentals, plus expertise in observability tooling such as OpenTelemetry, Prometheus, and Grafana.

- Working knowledge of relational databases (e.g., PostgreSQL, MySQL) , SQL, indexing, and query optimisation.

- Experience or familiarity with AI/ML concepts (e.g., LLMs, anomaly detection, or data-driven operations) applied to operational workflows.

- Excellent written and verbal communication skills, with the ability to brief executives during high-pressure incidents and influence senior technical stakeholders.

**Benefits:**

- Highly competitive salaries

- Comprehensive benefits package

- Learn more about NVIDIA benefits: www.nvidiabenefits.com

## Skills

### Required
- Site Reliability Engineering
- Production Engineering
- Platform Engineering
- Incident Management
- Distributed Systems
- Monitoring
- Reliability Engineering
- Programming Languages (Python, Go, Java)
- Public Cloud Platforms (AWS, Azure, GCP)
- Container Technologies (Docker, Kubernetes)
- Infrastructure-as-Code Tooling (Terraform, AWS CDK, CloudFormation)
- CI/CD Pipelines
- Linux/Unix
- Networking Fundamentals
- Observability Tooling (OpenTelemetry, Prometheus, Grafana)
- Relational Databases (PostgreSQL, MySQL)
- AI/ML Concepts (LLMs, Anomaly Detection)

### Nice to have
- Building & applying AI to operations
- Scaling reliability across distributed teams
- Ownership & community impact

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2024618?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
