# Senior Engineering Manager, Capacity Engineering

**Company**: Anthropic
**Location**: San Francisco, CA
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Salary**: $405,000-$485,000 USD
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q116758847

**Apply**: https://job-boards.greenhouse.io/anthropic/jobs/5363210008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_9ba1b38c-9ed

## Description

Anthropic manages one of the largest and fastest-growing infrastructure fleets in the industry , spanning multiple accelerator families, CPU families, and clouds. The Capacity Engineering team is responsible for making sure all of our infrastructure resources are accounted for, well-utilized, and efficiently allocated.

As the Senior Engineering Manager for Capacity Engineering, you will lead the team that builds and operates these production systems. You'll set technical direction, grow and develop a team of senior and staff-level engineers, and be accountable for the reliability and correctness of surfaces that leadership, research engineering, inference, infrastructure, and finance all depend on.

The team's work spans three overlapping areas:

- Data platform , Pipelines that ingest occupancy and utilization telemetry from Kubernetes clusters, normalize billing and usage across cloud providers, and serve the BigQuery tables the rest of the org queries against.

- Planning and Assurance , Making the state of the fleet legible and actionable in real time: cluster health tooling, capacity planning platforms, alerting on occupancy drops and allocation problems, and systemic fixes to scheduling and fragmentation.

- Efficiency , Measuring and improving how effectively every major workload uses the hardware it runs on, across training, inference, and evals.

Key Responsibilities:

- Be hands-on, lead and grow the team. Hire, onboard, coach, and retain senior and staff engineers.

- Champion your internal customers. Engage with them directly, bring what you learn back into the roadmap, and lead the team in building tools people genuinely want to use.

- Own the roadmap. Translate company-level compute strategy into a prioritized engineering roadmap across data platform, planning and efficiency.

- Set the technical bar. Review designs, weigh in on architecture, and hold the team to production standards.

- Run the team as a product organization. Ensure the team gathers its own requirements, defines schema contracts, and designs for a wide range of consumers.

- Be the primary partner for cross-functional stakeholders. Work closely with infrastructure, inference, research engineering, and finance leadership to align on capacity decisions, efficiency targets, and spend.

- Drive operational excellence. Own reliability and incident response for load-bearing systems, establish SLOs and on-call practices, and continuously reduce operational toil.

- Scale the function. As the fleet diversifies, anticipate where the team needs to grow in headcount, skills, and systems , and make the case for it.

What You Bring:

- Experience managing software or infrastructure engineering teams, including hiring senior engineers, managing performance, and developing people into larger scope.

- A strong technical background in production systems , data engineering, infrastructure, distributed systems, or observability.

- Familiarity with at least one major cloud provider (AWS, GCP, or Azure), Kubernetes-based infrastructure, and modern observability stacks.

- A track record of setting and executing an engineering roadmap in an ambiguous, high-autonomy environment with many stakeholders and shifting priorities.

- Excellent communication skills.

- Comfort owning operational responsibility for systems the company depends on, including on-call and incident management.

Preferred Qualifications:

- Experience leading teams working on capacity planning, resource management, product engineering or FinOps at a hyperscaler or in a large-scale ML environment.

- Familiarity with accelerator infrastructure , GPU metrics, TPU utilization, or ML training and inference systems at the hardware level.

- Experience with multi-cloud billing and telemetry normalization.

- Experience building or leading internal data products with self-service access, schema contracts, and documentation.

- Background in scheduling, packing efficiency, or profiling-driven optimization of large distributed workloads.

The annual compensation range for this role is $405,000-$485,000 USD.

## Skills

### Required
- data engineering
- infrastructure
- distributed systems
- observability
- cloud providers
- Kubernetes
- Python
- SQL

### Nice to have
- capacity planning
- resource management
- FinOps
- accelerator infrastructure
- multi-cloud billing
- telemetry normalization
- internal data products
- schema contracts
- documentation
- scheduling
- packing efficiency
- profiling-driven optimization

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/anthropic/jobs/5363210008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
