# Senior Staff Site Reliability Engineer

**Company**: NVIDIA
**Location**: Bengaluru
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2023198?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_1e89ecd7-0be

## Description

As a Senior Staff Site Reliability Engineer at NVIDIA, you will lead the technical strategy and roadmap for large-scale, multi-functional SRE initiatives that boost reliability, scalability, and developer efficiency throughout enterprise systems.

You will design and build resilient distributed systems that power NVIDIA's next-generation AI-powered enterprise products and services. This includes transforming legacy applications and database systems into modern and scalable architectures.

Your responsibilities will also involve:

- Driving automation and observability improvements using metrics and analytics, including AI workload quality signals and model performance telemetry.

- Building LLM-aware monitoring and autonomous incident response pipelines to reduce toil and accelerate MTTR.

- Collaborating with Cloud, Platform, Security, and AI/ML groups to develop modern SRE elements and AI-native platform features.

- Analyzing and running complex systems, including Kubernetes-scale and AI/ML infrastructure challenges.

- Driving AI-assisted and AI-first engineering practices across the organization.

To be successful in this role, you will need:

- 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.

- A BS degree in Computer Science or a related technical field.

- Strong proficiency in programming languages such as Python, Typescript, JavaScript, or Go.

- Experience with infrastructure-as-code tooling and observability implementations at scale.

- Deep expertise in systems architecture, networking, Kubernetes, and public cloud services.

- Outstanding problem-solving, communication, and collaboration skills.

NVIDIA offers highly competitive salaries and a comprehensive benefits package.

## Skills

### Required
- Site Reliability Engineering
- Cloud Architect
- Python
- Typescript
- JavaScript
- Go
- Kubernetes
- Public Cloud Services
- Observability
- Automation

### Nice to have
- Public Cloud
- Large-scale automation systems
- Agentic AI platforms
- LLM toolchains
- Agent orchestration frameworks

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2023198?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
