New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Site Reliability Engineer - Storage

NVIDIA
Apply →
senior full-time Santa Clara, CA

First indexed 2 Sept 2026

Description

NVIDIA is seeking a Senior Site Reliability Engineer to focus on HPC storage, playing a crucial role in designing, implementing, and optimizing on-prem High-Performance Computing (HPC) storage solutions while harnessing the power of cloud computing.

You will craft and deploy distributed storage solutions, build automation tools, and ensure the efficient operations of our growing IT ecosystem. Collaboration with engineering teams will be essential to align infrastructure with their evolving needs, document best practices, and contribute to groundbreaking projects.

Responsibilities:

  • Design, implement, and optimize on-prem HPC infrastructure supplemented with cloud computing.
  • Design and implement scalable and efficient Storage solutions for data-intensive applications.
  • Develop tooling to automate deployment, management, and monitoring of large-scale infrastructure environments.
  • Document general procedures and practices, perform technology evaluations related to distributed file systems.
  • Collaborate across teams to understand developers' workflows and gather infrastructure requirements.
  • Influence and guide methodologies for building, testing, and deploying applications.

Requirements:

  • BS in Computer Science (or equivalent experience) with 8+ years of relevant experience.
  • 8+ years of experience crafting technology solutions and resolving performance bottlenecks for HPC applications.
  • Experience with Enterprise NAS solutions like NetApp, Pure Storage, and S3-based storage.
  • Experience with parallel or distributed filesystems such as Lustre, GPFS.
  • Python/Bash/Golang programming/scripting experience.
  • Strong experience operating services in leading Cloud environments (AWS, Azure, GCP).
  • Experience with multiple monitoring stacks.
  • Excellent communication and collaboration skills.

Nice to Have:

  • Background with RDMA (InfiniBand or RoCE) fabrics.
  • Prior experience with HPC cluster management tools.
  • Experience with containerization technologies like Docker, Kubernetes.

NVIDIA offers equity and benefits, and is considered one of the technology world's most desirable employers.