Description
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline that designs, builds, and maintains large-scale production systems with high efficiency and availability.
You will lead the technical strategy and roadmap for large-scale, cross-functional SRE initiatives that improve reliability, scalability, and developer productivity across enterprise systems.
Responsibilities:
- Design, build, and maintain resilient distributed systems that power NVIDIA's next-generation AI-driven enterprise products and services.
- Architect and develop AI Agents, AI Skills to accelerate platform operations.
- Drive automation and observability improvements, using metrics and analytics to enhance performance, reliability, and efficiency.
- Collaborate across Cloud, Platform, Security, and AI/ML teams to implement modern SRE components that ensure high availability and secure operations.
- Analyze and troubleshoot complex systems, championing best practices in system design, incident management, and postmortem analysis.
- Mentor and influence engineers across teams, fostering technical excellence and a culture of reliability engineering.
Requirements:
- 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
- BS degree in Computer Science or a related technical field involving coding.
- Strong proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code.
- Experience with infrastructure-as-code such as AWS CDK, AWS CloudFormation, Terraform, or CrossPlane.
- Solid understanding of OpenTelemetry or other Observability implementation at scale.
- Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP).
- Outstanding problem-solving, communication, and teamwork skills, with the ability to influence across technical and interpersonal boundaries.
Benefits:
- Base salary range: 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.
- Equity and benefits package.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Staff-Site-Reliability-Engineer---AI-Platform-Runtime_JR2023969