New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Site Reliability Engineer - Storage

NVIDIA
Apply →
senior full-time Santa Clara, CA

First indexed 10 Aug 2026

Description

NVIDIA is seeking a Senior Site Reliability Engineer to focus on HPC storage, playing a crucial role in designing, implementing, and optimizing on-prem High-Performance Computing (HPC) storage solutions while harnessing the power of cloud computing.

You will craft and deploy distributed storage solutions, build automation tools, and ensure efficient operations of our growing IT ecosystem. Collaboration with engineering teams will be essential to align infrastructure with their evolving needs, document best practices, and contribute to groundbreaking projects.

Responsibilities:

  • Design and implement on-prem HPC infrastructure supplemented with cloud computing to support growing IT needs.
  • Design and implement scalable and efficient storage solutions for data-intensive applications, optimizing performance and cost-effectiveness.
  • Develop tooling to automate deployment, management, and operational monitoring of large-scale infrastructure environments.
  • Document general procedures and practices, perform technology evaluations related to distributed file systems.
  • Collaborate across teams to understand developers' workflows and gather infrastructure requirements.
  • Influence and guide methodologies for building, testing, and deploying applications to ensure optimal performance and resource utilization.

Requirements:

  • BS in Computer Science (or equivalent experience) with 8+ years of relevant experience, MS with 5+ years, or Ph.D. with 3 years.
  • 8+ years of experience crafting technology solutions and resolving performance bottlenecks for HPC applications.
  • Experience with enterprise NAS solutions like NetApp, Pure Storage, and S3-based storage such as Cloudian MinIO.
  • Experience with parallel or distributed filesystems like Lustre, GPFS.
  • Python/Bash/Golang programming/scripting experience.
  • Strong experience operating services in leading cloud environments (AWS, Azure, GCP).
  • Experience with multiple monitoring stacks like Prometheus+Grafana, Elasticsearch+Kibana, Splunk, Zabbix.
  • Excellent communication and collaboration skills.

Preferred Qualifications:

  • Background with RDMA (InfiniBand or RoCE) fabrics.
  • Prior experience with HPC cluster management tools like Slurm, PBS, LSF.
  • Experience with containerization technologies like Docker, Mesosphere DCOS, Kubernetes (k8s).

NVIDIA offers equity and benefits, and applications will be accepted until August 14, 2026.