Description
NVIDIA is seeking a Senior Site Reliability Engineer to work on the DGX Cloud team. The successful candidate will maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide.
Responsibilities:
- Build, implement, and support operational and reliability aspects of large-scale Kubernetes clusters with a focus on performance at scale, real-time monitoring, logging, and alerting.
- Define SLOs/SLIs, monitor error budgets, and streamline reporting.
- Support services before they launch through system creation consulting, developing software tools, platforms, and frameworks, capacity management, and launch reviews.
- Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
- Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.
- Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
- Lead triage and root-cause analysis of high-severity incidents.
- Practice balanced incident response and blameless postmortems.
- Participate in on-call rotation to support production services.
Requirements:
- BS in Computer Science or related technical field, or equivalent experience.
- 10+ years of experience operating production services.
- Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
- Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet).
- Proficiency in at least one high-level programming language (e.g., Python, Go).
- In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards.
- Proficient knowledge of SRE principles, encompassing SLOs, SLIs, error budgets, and incident handling.
- Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
Nice to have:
- Operating GPU-accelerated clusters with KubeVirt in production.
- Applying generative-AI techniques to reduce operational toil.
- Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Switzerland-Zurich/Senior-Site-Reliability-Engineer--DGX-Cloud_JR2021427-1