Description
NVIDIA is looking for a Senior AI/ML HPC Cluster Engineer to join our Managed AI Superclusters (MARS) team. You will provide leadership and strategic guidance on the management of large-scale HPC systems including the deployment of compute, networking, and storage.
Responsibilities:
- Provide leadership in systems administration and service delivery on our AI/HPC fleet by coordinating system upgrades, responding to incidents, and delivering reliability improvements.
- Collaborate closely with global teams to deliver a world-class user experience in AI and HPC research.
- Own day-to-day operations of production AI/HPC clusters, ensuring system health, user satisfaction, and efficient resource utilization.
- Develop and improve our ecosystem around GPU-accelerated computing, including developing scalable automation solutions.
- Build and maintain heterogeneous AI/ML clusters on-premises and in the cloud.
- Create and cultivate customer and cross-team relationships to meet user evolving user needs.
- Support our researchers to run their workloads, including performance analysis and optimizations.
- Analyze and optimize cluster efficiency, job fragmentation, and GPU waste to meet internal SLA targets.
- Conduct root cause analysis and suggest corrective action. Proactively find and fix issues before they occur.
- Lead SEV triage and postmortems for reliability incidents affecting users or infrastructure.
- Participate in on-call rotation and incident response for critical production GPU clusters.
Requirements:
- Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience.
- Minimum 5 years of experience designing and operating large-scale compute infrastructure.
- Experience with AI/HPC advanced job schedulers, such as Slurm, K8s, PBS, RTDA, BCM, or LSF.
- Proficient in administering Centos/RHEL and/or Ubuntu Linux distributions.
- Solid understanding of cluster configuration management tools (BCM, Terraform, Ansible, Puppet, Salt, etc.), container technologies (Docker, Singularity, Podman, Shifter, Charliecloud), Python programming, and bash scripting.
- Applied experience with AI/HPC workflows that use MPI.
- Experience analyzing and tuning performance for a variety of AI/HPC workloads.
- Passion for continual learning and staying ahead of emerging technologies and effective approaches in the HPC and AI/ML infrastructure fields.
Preferred Qualifications:
- Background with NVIDIA GPUs, CUDA Programming, NCCL and MLPerf benchmarking.
- Experience with AI/ML concepts, algorithms, models, and frameworks (PyTorch, Tensorflow).
- Experience with InfiniBand with IPoIB and RDMA.
- Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC workloads.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-HPC-Cluster-Engineer---AI--ML_JR2021817-1