New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Site Reliability Engineer

NVIDIA
Apply →
senior full-time 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333

First indexed 16 Sept 2026

Description

As a Senior Site Reliability Engineer at NVIDIA, you will work with the DGX Cloud team to maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide. Your role involves building, implementing, and supporting operational and reliability aspects of large-scale Kubernetes clusters, defining SLOs/SLIs, monitoring error allowances, and streamlining reporting.

Responsibilities:

  • Build, implement, and support operational and reliability aspects of large-scale Kubernetes clusters with a focus on performance at scale, real-time monitoring, logging, and alerting.
  • Define SLOs/SLIs, monitor error allowances, and streamline reporting.
  • Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews.
  • Maintain services once they are live by measuring and supervising availability, latency, and overall system health.
  • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.
  • Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
  • Lead triage and root-cause analysis of high-severity incidents.
  • Practice balanced incident response and blameless postmortems.
  • Participate in on-call rotation to support production services.

Requirements:

  • BS in Computer Science or related technical field, or equivalent experience.
  • 8+ years of experience operating production services.
  • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
  • Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet).
  • Proficiency in at least one high-level programming language (e.g., Python, Go).
  • In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards.
  • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.
  • Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.

Preferred qualifications:

  • Operating GPU-accelerated clusters with KubeVirt in production.
  • Applying generative-AI techniques to reduce operational toil.
  • Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.
  • Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis.

Benefits:

  • Base salary range: 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.
  • Equity and benefits package.