# Senior Staff Site Reliability Engineer

**Company**: NVIDIA
**Location**: Bengaluru
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Category**: IT
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2024324?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_9a5fce70-a7e

## Description

NVIDIA is seeking a highly skilled Senior Staff SRE to join its dynamic team. The company is at the forefront of technological innovation, driving efficiency and optimizing the performance of its infrastructure both on-prem and cloud.

**Job Summary:** Lead initiatives to transform IT Compute Core Team architecture to build new service offerings across On-Prem and Cloud. Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP, with a focus on performance, reliability, and automation.

**Key Responsibilities:**

- Lead initiatives to transform IT Compute Core Team architecture

- Design, scale, and deploy core infrastructure services

- Define and implement metrics to measure service efficiency

- Collect and review system data for capacity and planning purposes

- Develop and maintain tools for data collection, analysis, and visualization

- Collaborate with NVIDIA leadership and engineers to develop IT products and services

**Requirements:**

- Bachelor's degree in Engineering, Computer Science, Mathematics, or related field

- 15+ years of experience in compute platform engineering with a focus on automation

- Experience with containerization architectures and distributed systems infrastructure

- Strong analytical skills with the ability to define and track key performance metrics

- Proficiency in programming languages such as Go and/or Python

- Linux OS proficiency with kernel internals

**Nice to Have:**

- Deep understanding of infrastructure components like DNS, LDAP, and security tools

- Hands-on experience with containers and their implementation

- Deploying and managing services like DNS, LDAP at scale

- Solid understanding of microservices architecture, infrastructure as code (IaC), and configuration management tools

## Skills

### Required
- automation
- containerization
- distributed systems
- Linux
- Go
- Python
- Terraform
- Config Management

### Nice to have
- eBPF
- XDP
- observability
- DDoS mitigation
- microservices architecture
- infrastructure as code

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2024324?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
