# Staff Software Engineer, Node Infra

**Company**: Anthropic
**Location**: San Francisco, CA
**Work arrangement**: hybrid
**Experience**: staff
**Job type**: full-time
**Salary**: $320,000-$405,000 USD
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q116758847

**Apply**: https://job-boards.greenhouse.io/anthropic/jobs/5203868008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_67758c65-665

## Description

Anthropic's Infrastructure organization is foundational to its mission of developing AI systems that are reliable, interpretable, and steerable. The Node Infra team owns the full lifecycle of accelerator capacity at Anthropic, ingesting and provisioning compute from all major CSPs and datacenters, standing up and scaling clusters, and building health, diagnostics, and repair automation.

Key responsibilities:

- Own the technical strategy and roadmap for node lifecycle management

- Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families

- Design and operate systems that detect, isolate, and remediate unhealthy hardware automatically

- Define infrastructure architecture and ensure the hardest problems get solved

- Work closely with cloud providers and internal teams to shape long-term compute, data, and infrastructure strategy

- Establish and evolve operational excellence practices

- Support the growth of engineers through technical mentorship and coaching

Minimum qualifications:

- Deep expertise in distributed systems, reliability, and cloud platforms

- Strong proficiency in at least one systems language

- Hands-on experience with machine learning accelerators

- Track record of leading complex technical initiatives

- Ability to build alignment across senior stakeholders and communicate effectively

Preferred qualifications:

- 8+ years of software engineering experience

- Experience managing large-scale compute infrastructure

- Depth in Kubernetes internals, cluster orchestration systems, or node provisioning pipelines

- Low-level systems experience

- Familiarity with high-performance networking

- Demonstrated ownership of production reliability

- Contributions to relevant open-source projects

The annual compensation range for this role is $320,000-$405,000 USD.

## Skills

### Required
- distributed systems
- cloud platforms
- systems languages
- machine learning accelerators

### Nice to have
- Kubernetes
- cluster orchestration
- node provisioning
- high-performance networking
- open-source projects

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/anthropic/jobs/5203868008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
