Description
NVIDIA is seeking a highly skilled Senior Staff SRE to join its dynamic team. The company is at the forefront of technological innovation, driving efficiency and optimizing the performance of its infrastructure both on-prem and cloud.
Job Summary: Lead initiatives to transform IT Compute Core Team architecture to build new service offerings across On-Prem and Cloud. Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP, with a focus on performance, reliability, and automation.
Key Responsibilities:
- Lead initiatives to transform IT Compute Core Team architecture
- Design, scale, and deploy core infrastructure services
- Define and implement metrics to measure service efficiency
- Collect and review system data for capacity and planning purposes
- Develop and maintain tools for data collection, analysis, and visualization
- Collaborate with NVIDIA leadership and engineers to develop IT products and services
Requirements:
- Bachelor's degree in Engineering, Computer Science, Mathematics, or related field
- 15+ years of experience in compute platform engineering with a focus on automation
- Experience with containerization architectures and distributed systems infrastructure
- Strong analytical skills with the ability to define and track key performance metrics
- Proficiency in programming languages such as Go and/or Python
- Linux OS proficiency with kernel internals
Nice to Have:
- Deep understanding of infrastructure components like DNS, LDAP, and security tools
- Hands-on experience with containers and their implementation
- Deploying and managing services like DNS, LDAP at scale
- Solid understanding of microservices architecture, infrastructure as code (IaC), and configuration management tools
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Staff-Site-Reliability-Engineer_JR2024324