New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior HPC Cluster Engineer - AI, ML

NVIDIA
Apply →
senior full-time

First indexed 6 Aug 2026

Description

NVIDIA is looking for a Senior AI/ML HPC Cluster Engineer to join our Managed AI Superclusters (MARS) team. You will provide leadership and strategic guidance on the management of large-scale HPC systems including the deployment of compute, networking, and storage.

Responsibilities:

  • Provide leadership in systems administration and service delivery on our AI/HPC fleet by coordinating system upgrades, responding to incidents, and delivering reliability improvements.
  • Collaborate closely with global teams to deliver a world-class user experience in AI and HPC research.
  • Own day-to-day operations of production AI/HPC clusters, ensuring system health, user satisfaction, and efficient resource utilization.
  • Develop and improve our ecosystem around GPU-accelerated computing, including developing scalable automation solutions.
  • Build and maintain heterogeneous AI/ML clusters on-premises and in the cloud.
  • Create and cultivate customer and cross-team relationships to meet user evolving user needs.
  • Support our researchers to run their workloads, including performance analysis and optimizations.
  • Analyze and optimize cluster efficiency, job fragmentation, and GPU waste to meet internal SLA targets.
  • Conduct root cause analysis and suggest corrective action. Proactively find and fix issues before they occur.
  • Lead SEV triage and postmortems for reliability incidents affecting users or infrastructure.
  • Participate in on-call rotation and incident response for critical production GPU clusters.

Requirements:

  • Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience.
  • Minimum 5 years of experience designing and operating large-scale compute infrastructure.
  • Experience with AI/HPC advanced job schedulers, such as Slurm, K8s, PBS, RTDA, BCM, or LSF.
  • Proficient in administering Centos/RHEL and/or Ubuntu Linux distributions.
  • Solid understanding of cluster configuration management tools (BCM, Terraform, Ansible, Puppet, Salt, etc.), container technologies (Docker, Singularity, Podman, Shifter, Charliecloud), Python programming, and bash scripting.
  • Applied experience with AI/HPC workflows that use MPI.
  • Experience analyzing and tuning performance for a variety of AI/HPC workloads.
  • Passion for continual learning and staying ahead of emerging technologies and effective approaches in the HPC and AI/ML infrastructure fields.

Preferred Qualifications:

  • Background with NVIDIA GPUs, CUDA Programming, NCCL and MLPerf benchmarking.
  • Experience with AI/ML concepts, algorithms, models, and frameworks (PyTorch, Tensorflow).
  • Experience with InfiniBand with IPoIB and RDMA.
  • Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC workloads.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-HPC-Cluster-Engineer---AI--ML_JR2021817-1