Description
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology.
We are looking for a Senior AI Infrastructure Engineer to lead the vision, execution, and long-term stability of how Anduril trains with GPUs at scale.
Responsibilities
- Design, build, and operate large-scale GPU computing infrastructure for machine learning and AI applications.
- Ensure the robustness, availability, and fault tolerance of high-performance GPU systems.
- Develop and implement automated deployment and configuration tools for infrastructure.
- Collaborate with cross-functional teams to understand emerging compute needs and translate them into platform capabilities.
- Monitor, troubleshoot, and optimize the performance of GPU clusters and storage systems.
- Provide technical support and escalation for engineers and researchers using the platform.
Requirements
- 10+ years of experience in hands-on infrastructure, HPC, or datacenter engineering roles supporting GPU compute at scale.
- Hands-on experience with H200/B200/B300 GPU systems, high-performance interconnects, and parallel storage systems.
- Strong background in automation, Kubernetes, and GPU scheduling.
- Ability to lift/move 50+ lbs and perform physical datacenter work.
- Eligible to obtain and maintain an active U.S. Top Secret clearance.
Preferred Qualifications
- Experience with NVIDIA NVL72 rack-scale systems.
- Knowledge of network fabric tuning and congestion control.
- Familiarity with GPU/network observability tooling and automated fault detection.
Benefits
Anduril offers a comprehensive benefits package, including competitive salary, equity grants, and top-tier benefits for full-time employees.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/andurilindustries/jobs/5227989007