New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Site Reliability Engineer

NVIDIA
Apply →
senior full-time Israel, Yokneam

First indexed 26 Aug 2026

Description

We are seeking a highly skilled Principal Staff SRE to join our dynamic Core Infrastructure team. This role will lead initiatives that improve the efficiency, performance, reliability, and evolution of infrastructure services across on-premises and cloud environments. Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package.

What you'll be doing:

  • Lead initiatives that transform the IT Compute Core architecture and build new infrastructure service offerings across on-premises and cloud environments.
  • Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP, with responsibility for performance, reliability, automation, monitoring, high availability, capacity planning, and lifecycle management at global scale.
  • Define and implement service-efficiency metrics, and drive improvements through software and hardware optimization, including SR-IOV and DPU capabilities where appropriate.
  • Apply technologies such as eBPF and XDP to improve observability, performance analysis, and DDoS-mitigation capabilities.
  • Collect and analyze system and capacity data; develop enterprise-wide capacity plans and coordinate with management to implement appropriate changes.
  • Develop and maintain tools for collecting, analyzing, and visualizing data for reporting, alerting, and monitoring.
  • Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop compelling IT products and services that meet customer needs.

What we need to see:

  • Bachelor's degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
  • 12+ years of demonstrable experience in compute platform engineering, with a strong focus on automation and technical leadership in large-scale environments.
  • Experience designing and deploying containerization architectures and distributed-systems infrastructure.
  • Demonstrable experience evaluating existing application architectures and seeing opportunities for containerization that improve scalability, reliability, and efficiency.
  • Strong analytical skills, including the ability to define and track key performance metrics.
  • Experience developing tools for data analysis and performance profiling, including Terraform and configuration-management tools.
  • Proficiency in Go, Python, or similar programming languages.
  • Linux OS proficiency, including kernel internals.
  • Experience operating large-scale environments with bare-metal build infrastructure.
  • Understanding of network protocols and architectures, including VLAN, VXLAN, SDN, BGP, and Anycast.

Ways to stand out from the crowd:

  • Deep understanding of complementary infrastructure components, including DNS, LDAP, and security tools.
  • Hands-on experience with containers and their implementation.
  • Experience deploying and managing services such as DNS and LDAP at scale.
  • Solid understanding of microservices architecture, infrastructure as code, and configuration-management tools.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Israel-Yokneam/Senior-Site-Reliability-Engineer_JR2024251