New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Site Reliability Engineer - Cloud

NVIDIA
Apply →
senior full-time Santa Clara, CA

First indexed 4 Aug 2026

Description

NVIDIA's Digital Marketing Organisation seeks a senior Site Reliability Engineer (SRE) to join the Santa Clara, CA team. As an SRE at NVIDIA, you will have a meaningful role in keeping Digital Marketing Services reliable, fast, and efficient.

Responsibilities:

  • Build and deploy large-scale dynamic URL redirects using Akamai Edge Redirector Cloudlets for promotional efforts and site migrations.
  • Configure Akamai Forward Rewrite Cloudlets to map inbound requests to SEO-friendly paths.
  • Provide on-call support for production-grade applications, responding to incidents, prioritizing issues, and driving resolution across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.
  • Author, test, and activate shared and non-shared Cloudlet Policies via the Akamai Cloudlets Policy Manager.
  • Maintain custom match criteria , including Geo, Device Characteristics, RegEx, and Query Strings , to ensure efficient origin offload and intelligent content delivery.
  • Quickly identify and address user-reported problems throughout the Digital Marketing Organisation ecosystem.
  • On-board new applications, AI/ML services, and model endpoints on AWS Infrastructure.
  • Implement monitors, alerts, and SOPs to ensure early detection and accurate response to service-impacting issues, including tracking model drift and inference latency.

Requirements:

  • MS or BS in Computer Science/Engineering or a related field, or equivalent experience.
  • 8+ years’ experience supporting technical operations in a live-site production environment with a real passion for CDN automation, tooling, and infrastructure supporting AI applications.
  • Strong knowledge of the Kubernetes Platform, deployments, and cloud-native automation.
  • Proven strengths in problem-solving and root-causing issues, while continuously seeking ways to drive optimization, efficiency, and the bottom line.
  • Advanced level experience with scripting and development in Python, fully automating operational steps with “one-click” rapid solutions.
  • Key participation in the incident management process for early recognition of all service-impacting issues, accurate triage, partner communication, impact containment, service restoration, and post-incident follow-up. SRE On call experience is a must.

Benefits:

  • Highly competitive salaries
  • Comprehensive benefits package
  • Equity eligibility
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Site-Reliability-Engineer---Cloud_JR2022489