# Senior Site Reliability Engineer

**Company**: NVIDIA
**Experience**: senior
**Job type**: full-time
**Salary**: 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Site-Reliability-Engineer--DGX-Cloud_JR2025495?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_a13ed822-2f6

## Description

As a Senior Site Reliability Engineer at NVIDIA, you will work with the DGX Cloud team to maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide. Your role involves building, implementing, and supporting operational and reliability aspects of large-scale Kubernetes clusters, defining SLOs/SLIs, monitoring error allowances, and streamlining reporting.

Responsibilities:

- Build, implement, and support operational and reliability aspects of large-scale Kubernetes clusters with a focus on performance at scale, real-time monitoring, logging, and alerting.

- Define SLOs/SLIs, monitor error allowances, and streamline reporting.

- Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews.

- Maintain services once they are live by measuring and supervising availability, latency, and overall system health.

- Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.

- Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.

- Lead triage and root-cause analysis of high-severity incidents.

- Practice balanced incident response and blameless postmortems.

- Participate in on-call rotation to support production services.

Requirements:

- BS in Computer Science or related technical field, or equivalent experience.

- 8+ years of experience operating production services.

- Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.

- Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet).

- Proficiency in at least one high-level programming language (e.g., Python, Go).

- In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards.

- Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.

- Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.

Preferred qualifications:

- Operating GPU-accelerated clusters with KubeVirt in production.

- Applying generative-AI techniques to reduce operational toil.

- Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.

- Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis.

Benefits:

- Base salary range: 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.

- Equity and benefits package.

## Skills

### Required
- Kubernetes
- containerization
- microservices architecture
- Terraform
- Ansible
- Chef
- Puppet
- Python
- Go
- Linux
- TCP/IP
- cloud security standards
- SRE principles
- OpenTelemetry
- Prometheus
- Grafana
- ELK Stack
- Lightstep
- Splunk

### Nice to have
- GPU-accelerated clusters
- KubeVirt
- generative-AI techniques
- workflow orchestration platforms
- Temporal
- Cadence
- Airflow
- Argo Workflows
- Step Functions
- vLLM
- SGLang
- PyTorch
- TensorRT-LLM
- NVIDIA Dynamo
- CUDA
- NCCL
- GPU performance analysis

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Site-Reliability-Engineer--DGX-Cloud_JR2025495?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
