Description
We are seeking a highly skilled Principal Staff SRE to join our dynamic Core Infrastructure team. This role will lead initiatives that improve the efficiency, performance, reliability, and evolution of infrastructure services across on-premises and cloud environments. Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package.
What you'll be doing:
- Lead initiatives that transform the IT Compute Core architecture and build new infrastructure service offerings across on-premises and cloud environments.
- Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP, with responsibility for performance, reliability, automation, monitoring, high availability, capacity planning, and lifecycle management at global scale.
- Define and implement service-efficiency metrics, and drive improvements through software and hardware optimization, including SR-IOV and DPU capabilities where appropriate.
- Apply technologies such as eBPF and XDP to improve observability, performance analysis, and DDoS-mitigation capabilities.
- Collect and analyze system and capacity data; develop enterprise-wide capacity plans and coordinate with management to implement appropriate changes.
- Develop and maintain tools for collecting, analyzing, and visualizing data for reporting, alerting, and monitoring.
- Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop compelling IT products and services that meet customer needs.
What we need to see:
- Bachelor's degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
- 12+ years of demonstrable experience in compute platform engineering, with a strong focus on automation and technical leadership in large-scale environments.
- Experience designing and deploying containerization architectures and distributed-systems infrastructure.
- Demonstrable experience evaluating existing application architectures and seeing opportunities for containerization that improve scalability, reliability, and efficiency.
- Strong analytical skills, including the ability to define and track key performance metrics.
- Experience developing tools for data analysis and performance profiling, including Terraform and configuration-management tools.
- Proficiency in Go, Python, or similar programming languages.
- Linux OS proficiency, including kernel internals.
- Experience operating large-scale environments with bare-metal build infrastructure.
- Understanding of network protocols and architectures, including VLAN, VXLAN, SDN, BGP, and Anycast.
Ways to stand out from the crowd:
- Deep understanding of complementary infrastructure components, including DNS, LDAP, and security tools.
- Hands-on experience with containers and their implementation.
- Experience deploying and managing services such as DNS and LDAP at scale.
- Solid understanding of microservices architecture, infrastructure as code, and configuration-management tools.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Israel-Yokneam/Senior-Site-Reliability-Engineer_JR2024251