# Senior Site Reliability Engineer, DGX Cloud

**Company**: NVIDIA
**Location**: Zurich
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Switzerland-Zurich/Senior-Site-Reliability-Engineer--DGX-Cloud_JR2021427-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_8e923b92-1ad

## Description

NVIDIA is seeking a Senior Site Reliability Engineer to work on the DGX Cloud team. The successful candidate will maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide.

Responsibilities:

- Build, implement, and support operational and reliability aspects of large-scale Kubernetes clusters with a focus on performance at scale, real-time monitoring, logging, and alerting.

- Define SLOs/SLIs, monitor error budgets, and streamline reporting.

- Support services before they launch through system creation consulting, developing software tools, platforms, and frameworks, capacity management, and launch reviews.

- Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.

- Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.

- Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.

- Lead triage and root-cause analysis of high-severity incidents.

- Practice balanced incident response and blameless postmortems.

- Participate in on-call rotation to support production services.

Requirements:

- BS in Computer Science or related technical field, or equivalent experience.

- 10+ years of experience operating production services.

- Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.

- Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet).

- Proficiency in at least one high-level programming language (e.g., Python, Go).

- In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards.

- Proficient knowledge of SRE principles, encompassing SLOs, SLIs, error budgets, and incident handling.

- Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.

Nice to have:

- Operating GPU-accelerated clusters with KubeVirt in production.

- Applying generative-AI techniques to reduce operational toil.

- Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.

## Skills

### Required
- Kubernetes administration
- containerization
- microservices architecture
- infrastructure automation tools
- high-level programming languages
- Linux operating systems
- networking fundamentals
- cloud security standards
- SRE principles
- observability stacks

### Nice to have
- GPU-accelerated clusters with KubeVirt
- generative-AI techniques
- workflow orchestration platforms

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Switzerland-Zurich/Senior-Site-Reliability-Engineer--DGX-Cloud_JR2021427-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
