Description
Job Overview
We're looking for a Staff Infrastructure Engineer to join our Cluster Infrastructure team at Anthropic. As a Staff engineer on this team, you'll set the technical direction for how Anthropic brings compute online - at a moment when the scale of that compute is growing faster than at almost any company in the world.
Responsibilities
- Own the technical strategy and roadmap for agent-driven cluster lifecycle management - provisioning, updates and decommissioning
- Partner across teams to ensure new compute capacity is ingested on time
- Align with partner teams on physical build-out and leverage cloud solutions to deliver high-bandwidth inter-cluster connectivity
- Collaborate with security owners to ensure clusters are provisioned secure-by-default
- Define and drive strategy on cluster scalability, homogeneity and fault tolerance
- Work closely with cloud providers and internal research, inference and product teams to shape long-term compute, data, and infrastructure strategy
- Establish and evolve operational-excellence practices: incident response, postmortem culture and on-call health
- Support the growth of engineers around you through technical mentorship and coaching
Requirements
- Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure)
- Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform
- Track record of leading complex, multi-quarter technical initiatives spanning multiple teams or systems
- Ability to build alignment across senior stakeholders and communicate effectively at all levels
Preferred Qualifications
- 8+ years of software engineering experience, including time as a technical lead setting direction for a team
- Experience operating large-scale compute infrastructure at hyperscale (100+ clusters, 10K+ nodes)
- Depth in one or more of: Kubernetes internals, cluster provisioning and management systems, cluster orchestration systems (Mesos, Borg-like)
- Experience with cloud networking: VPC design and peering, Shared VPC/Transit Gateway, Cloud Interconnect/Direct Connect, Cloud NAT, cross-cloud private connectivity, BGP and route control, edge load balancing and DDoS mitigation (Cloud Armor / AWS Shield)
- Experience with cluster and host networking: CNI (Cilium), eBPF, NetworkPolicy, multi-NIC, sFlow, service mesh (Istio/Envoy/Linkerd, mTLS)
- Experience with cluster security: pod security standards and admission control, RBAC and least-privilege IAM, node and container hardening, supply-chain/image provenance
- Deep experience with infrastructure-as-code (Terraform, Atlantis), workflow orchestration (Temporal, Argo Workflows)
- Skill in quickly understanding systems design tradeoffs and keeping track of rapidly evolving software systems
Compensation
The annual compensation range for this role is $320,000-$4,050,000 USD.
Logistics
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
- Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.
- Visa sponsorship: We do sponsor visas!
Benefits
We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/anthropic/jobs/5206978008