Description
NVIDIA is looking for a Senior Staff Site Reliability Engineer to join our team in India. This is a senior individual-contributor role that combines real-time incident leadership with deep hands-on engineering, focused on how we detect, respond to, and prevent issues at scale across NVIDIA's AI-powered enterprise platforms.
You will operate as an Incident Commander during critical events, set the technical direction for reliability engineering across multiple teams, and build the automation, observability, and AI-assisted tooling that reduce operational toil and raise reliability over time. You will also mentor and grow SRE talent as we scale the practice in India.
Responsibilities:
- Lead major incidents end to end , driving triage, cross-team coordination, decision making, and executive communication across global time zones.
- Set technical direction for SRE initiatives that improve reliability, scalability, and developer efficiency across NVIDIA's enterprise systems, and drive them to adoption beyond your immediate team.
- Design, build, and operate distributed systems , including Kubernetes-based and cloud-native infrastructure , that power NVIDIA's AI-powered enterprise products and services.
- Build automation for incident detection, triage, communication, and remediation, replacing manual runbooks with self-healing systems.
- Improve observability and signal quality to enable earlier detection, reduce alert noise, and eliminate reliance on user-reported issues.
- Drive root cause analysis and translate learnings into systemic fixes, automation, and prevention mechanisms; raise the bar on post-incident review quality across the org.
- Apply AI and data-driven techniques , LLMs, anomaly detection, signal correlation , to enhance incident triage, summarisation, and decision support.
- Champion AI-assisted engineering practices, including coding agents and LLM-powered tooling, to accelerate day-to-day engineering workflows.
- Partner with Cloud, Platform, Security, and AI/ML teams to embed SRE best practices, define SLOs and error budgets, and influence architecture early in the design cycle.
- Mentor engineers, raise engineering standards through design and code review, and help build a strong reliability culture in the India organisation.
Requirements:
- 10+ years of experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management roles, with a track record of technical leadership at scale.
- BS or MS degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
- Proven experience acting as an Incident Commander or leading major incident response in complex, high-availability environments.
- Deep understanding of distributed systems, monitoring, and reliability engineering principles , SLIs/SLOs, error budgets, capacity planning, and graceful degradation.
- Strong proficiency in at least one programming language (e.g., Python, Go, Java ) to build production-grade automation and tooling.
- Hands-on expertise with public cloud platforms (AWS, Azure, or GCP) and container technologies such as Docker and Kubernetes.
- Solid experience with infrastructure-as-code tooling (e.g., Terraform, AWS CDK, CloudFormation) and CI/CD pipelines.
- Strong Linux/Unix and networking fundamentals, plus expertise in observability tooling such as OpenTelemetry, Prometheus, and Grafana.
- Working knowledge of relational databases (e.g., PostgreSQL, MySQL) , SQL, indexing, and query optimisation.
- Experience or familiarity with AI/ML concepts (e.g., LLMs, anomaly detection, or data-driven operations) applied to operational workflows.
- Excellent written and verbal communication skills, with the ability to brief executives during high-pressure incidents and influence senior technical stakeholders.
Benefits:
- Highly competitive salaries
- Comprehensive benefits package
- Learn more about NVIDIA benefits: www.nvidiabenefits.com
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2024618