Description
Anthropic's Infrastructure organization is foundational to its mission of developing AI systems that are reliable, interpretable, and steerable. The Node Infra team owns the full lifecycle of accelerator capacity at Anthropic, ingesting and provisioning compute from all major CSPs and datacenters, standing up and scaling clusters, and building health, diagnostics, and repair automation.
Key responsibilities:
- Own the technical strategy and roadmap for node lifecycle management
- Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families
- Design and operate systems that detect, isolate, and remediate unhealthy hardware automatically
- Define infrastructure architecture and ensure the hardest problems get solved
- Work closely with cloud providers and internal teams to shape long-term compute, data, and infrastructure strategy
- Establish and evolve operational excellence practices
- Support the growth of engineers through technical mentorship and coaching
Minimum qualifications:
- Deep expertise in distributed systems, reliability, and cloud platforms
- Strong proficiency in at least one systems language
- Hands-on experience with machine learning accelerators
- Track record of leading complex technical initiatives
- Ability to build alignment across senior stakeholders and communicate effectively
Preferred qualifications:
- 8+ years of software engineering experience
- Experience managing large-scale compute infrastructure
- Depth in Kubernetes internals, cluster orchestration systems, or node provisioning pipelines
- Low-level systems experience
- Familiarity with high-performance networking
- Demonstrated ownership of production reliability
- Contributions to relevant open-source projects
The annual compensation range for this role is $320,000-$405,000 USD.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/anthropic/jobs/5203868008