New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior GPU and HPC Infrastructure Engineer - DGX Cloud

NVIDIA
Apply →
remote senior full-time Santa Clara, CA

First indexed 8 Jul 2026

Description

NVIDIA is seeking a Senior GPU and HPC Infrastructure Engineer to scale its AI Infrastructure. You will contribute to building and deploying leading infrastructure solutions for AI-based applications.

Job Summary: As a Senior GPU and HPC Infrastructure Engineer, you will work on automating GPU asset provisioning, configuration, and lifecycle management across cloud providers. You will implement monitoring and health management capabilities, work on software managing NVLINK topography across GPU clusters, and build automated test infrastructure.

Responsibilities:

  • Contribute to the platform that automates GPU asset provisioning, configuration, and lifecycle management across cloud providers.
  • Implement monitoring and health management capabilities for industry-leading reliability, availability, and scalability of GPU assets.
  • Work on software that manages NVLINK topography across GPU clusters.
  • Build automated test infrastructure to qualify distributed systems for operation.
  • Collaborate with engineering teams across NVIDIA to ensure seamless integration from hardware to AI training applications.

Requirements:

  • 10+ years of software engineering experience on large-scale production systems.
  • BS in Computer Science/Engineering/Physics/Mathematics or equivalent experience.
  • Expert-level knowledge of a systems programming language (Go, Python) and solid understanding of Data Structure and Algorithms.
  • Expert-level knowledge of Linux system administration and management.
  • Understanding of cluster management systems (Kubernetes, SLURM).
  • Understanding of performance, security, and reliability in complex distributed systems.

Preferred Qualifications:

  • Proficiency in architecting and managing large-scale distributed systems, independent of cloud providers.
  • Deep knowledge of datacenter operations and GPU hardware.
  • Hands-on experience working with RDMA networking.
  • Advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, SLURM).
  • Hands-on experience in Machine Learning Operations.
  • Hands-on experience with Bright Cluster Manager.

Benefits:

  • Equity and benefits package