Description
ElevenLabs is seeking an HPC Infrastructure Engineer to join its research infrastructure team. The successful candidate will operate and improve NVIDIA GPU clusters across bare metal and rented capacity.
Responsibilities:
- Operate and improve the GPU fleet end-to-end: provisioning, scheduling, monitoring, upgrades, capacity planning
- Build automation that keeps the fleet healthy without human intervention , node health checks, automated draining and remediation, burn-in pipelines for new capacity
- Own the stack beneath the training code: OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, high-speed networking (InfiniBand/RoCE)
- Run and tune job scheduling (Slurm or similar) so researchers get compute fairly and fast
- Build and maintain high-performance storage for datasets and checkpoints
- Hunt down performance problems: stragglers, degraded links, thermal issues, flaky GPUs , and fix the class of problem, not just the instance
- Evaluate rented GPU capacity: benchmark it, validate it, hold providers to their SLAs
- Hands-on hardware work when it's needed: racking, cabling, diagnostics, coordinating with datacenter staff and vendors
- Keep clusters secure by default: access control, network isolation, secrets
Requirements:
- Experience running large-scale Linux server or GPU environments in production
- Knowledge of the NVIDIA stack: drivers, CUDA, NCCL, DCGM
- Comfortable with bare-metal environments, server hardware, and high-speed networking
- Proficiency in Python and/or Bash, with IaC tools like Ansible or Terraform
- Experience with metrics, logs, and PromQL
Nice to have:
- Experience supporting ML training workloads from the infra side
- Experience evaluating and working with GPU cloud providers
- Knowledge of parallel filesystems or large-scale object storage
- BMC/IPMI/Redfish automation, PXE provisioning at scale
- Power and cooling awareness for dense GPU deployments
Location: Remote
Benefits:
- Innovative culture
- Growth paths
- Learning & development support
- Social travel stipend
- Annual company offsite
- Co-working stipend
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://elevenlabs.io/sv/careers/120da2b3-d88b-4e3c-9b89-d19ff73db9d9/hpc-infrastructure-engineer-gpu-clusters