# Senior SRE Engineer

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-SRE-Engineer_JR2020072?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_dff2a5b7-0b4

## Description

NVIDIA is seeking a Senior SRE Engineer to join its Infrastructure, Planning, and Processes organization. You will own and scale the internal CI as a Service platform, managing it like a product to ensure high availability, self-service, observability, and elasticity.

The platform includes the shared GitLab CI and GitHub Actions infrastructure used daily by thousands of engineers. The team collaborates with other NVIDIA Software units to meet their infrastructure and system needs.

**Responsibilities:**

- Develop, handle, and expand a multi-tenant CI platform built on GitLab's CI framework and GitHub's action-based automation.

- Own the underlying Kubernetes substrate end-to-end, including cluster lifecycle, upgrades, and autoscaling.

- Drive reliability and capacity engineering, including SLOs and error budgets for queue time, job success, and runner availability.

- Build the self-service layer pipeline templates, reusable workflows, golden images, policy-as-code, and guardrails.

- Improve developer experience continuously, focusing on faster cold-starts, smarter caching, and deep observability into pipeline performance and cost.

**Requirements:**

- 5+ years in SRE/platform roles with strong fundamentals in SLO/SLI build, incident command, resource planning, and production Linux administration at scale.

- Deep Kubernetes administration experience, including CRDs and operators, HPA/VPA/cluster-autoscaling, and service mesh.

- Hands-on expertise with GitLab continuous integration and GitHub automated workflows at scale.

- Strong scripting and automation skills in Python, Go, bash scripting, or equivalent.

- BS/MS in CS or equivalent experience in building observability tools.

**Preferred Qualifications:**

- Strong understanding of containerization and microservices architecture.

- Certified Kubernetes Administrator (CKA), Certified Kubernetes Security Specialist (CKS) & Certified Kubernetes Application Developer (CKAD) preferred.

- Built or extended the CI control plane itself, including custom runner schedulers, autoscaling, and pipeline orchestration.

You will also be eligible for equity and benefits.

## Skills

### Required
- Kubernetes
- GitLab CI
- GitHub Actions
- Python
- Go
- bash scripting
- Linux administration
- SLO/SLI
- incident command
- resource planning
- production Linux administration
- containerization
- microservices architecture

### Nice to have
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Security Specialist (CKS)
- Certified Kubernetes Application Developer (CKAD)
- custom runner schedulers
- autoscaling
- pipeline orchestration

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-SRE-Engineer_JR2020072?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
