New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior HPC AI Cluster Engineer

NVIDIA
Apply →
remote senior full-time Santa Clara, CA

First indexed 21 Aug 2026

Description

NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. The team focuses on building supercomputers and AI clusters based on groundbreaking technologies. As a Senior HPC AI Cluster Engineer, you will be a key player in designing and implementing large-scale HPC/AI clusters, providing insights on at-scale system design and tuning mechanisms for large-scale compute runs.

Responsibilities:

  • Design, implement, and maintain large-scale HPC/AI clusters with monitoring, logging, and alerting
  • Manage Linux job/workload schedules and orchestration tools
  • Develop and maintain continuous integration and delivery pipelines
  • Develop tooling to automate deployment and management of large-scale infrastructure environments, automate operational monitoring and alerting, and enable self-service consumption of resources
  • Deploy monitoring solutions for servers, network, and storage
  • Perform troubleshooting from bare metal, operating system, software stack, and application level
  • Develop, redefine, and document standard methodologies to share with internal teams
  • Support Research & Development activities and engage in POCs/POVs for future improvements

Requirements:

  • A degree in Computer Science, Engineering, or a related field (or equivalent experience) and 8+ years of experience
  • Knowledge of HPC and AI solution technologies from CPUs and GPUs to high-speed interconnects and supporting software
  • Experience with job scheduling workloads and orchestration tools such as Slurm, K8s
  • Excellent knowledge of Windows and Linux (Redhat/CentOS and Ubuntu) networking and internals, ACLs, and OS-level security protection and common protocols
  • Experience with multiple storage solutions such as Lustre, GPFS, Weka.io
  • Python programming and bash scripting experience
  • Comfortable with automation and configuration management tools such as Jenkins, Ansible, Puppet/Chef
  • Deep knowledge of Networking Protocols like InfiniBand, Ethernet
  • Deep understanding and experience with virtual systems
  • Familiarity with cloud computing platforms

Preferred Qualifications:

  • Knowledge of CPU and/or GPU architecture
  • Knowledge of Kubernetes, container-related microservice technologies
  • Experience with GPU-focused hardware/software (DGX, Cuda)
  • Experience with RDMA (InfiniBand or RoCE) fabrics
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-HPC-AI-Cluster-Engineer_JR2023577