Description
As a Senior Staff Site Reliability Engineer at NVIDIA, you will lead the technical strategy and roadmap for large-scale, multi-functional SRE initiatives that boost reliability, scalability, and developer efficiency throughout enterprise systems.
You will design and build resilient distributed systems that power NVIDIA's next-generation AI-powered enterprise products and services. This includes transforming legacy applications and database systems into modern and scalable architectures.
Your responsibilities will also involve:
- Driving automation and observability improvements using metrics and analytics, including AI workload quality signals and model performance telemetry.
- Building LLM-aware monitoring and autonomous incident response pipelines to reduce toil and accelerate MTTR.
- Collaborating with Cloud, Platform, Security, and AI/ML groups to develop modern SRE elements and AI-native platform features.
- Analyzing and running complex systems, including Kubernetes-scale and AI/ML infrastructure challenges.
- Driving AI-assisted and AI-first engineering practices across the organization.
To be successful in this role, you will need:
- 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
- A BS degree in Computer Science or a related technical field.
- Strong proficiency in programming languages such as Python, Typescript, JavaScript, or Go.
- Experience with infrastructure-as-code tooling and observability implementations at scale.
- Deep expertise in systems architecture, networking, Kubernetes, and public cloud services.
- Outstanding problem-solving, communication, and collaboration skills.
NVIDIA offers highly competitive salaries and a comprehensive benefits package.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2023198