New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Anthropic

Senior Engineering Manager, Capacity Engineering

Anthropic
Apply →
hybrid senior full-time $405,000-$485,000 USD San Francisco, CA

First indexed 12 Sept 2026

Description

Anthropic's mission is to create reliable, interpretable, and steerable AI systems. The Capacity Engineering team is responsible for ensuring all infrastructure resources are accounted for, well-utilized, and efficiently allocated.

As the Senior Engineering Manager for Capacity Engineering, you will lead the team that builds and operates these production systems. You will set technical direction, grow and develop a team of senior and staff-level engineers, and be accountable for the reliability and correctness of surfaces that leadership, research engineering, inference, infrastructure, and finance all depend on.

The team's work spans three overlapping areas:

  • Data platform: pipelines that ingest occupancy and utilization telemetry from Kubernetes clusters, normalize billing and usage across cloud providers, and serve BigQuery tables.
  • Planning and Assurance: making the state of the fleet legible and actionable in real time, cluster health tooling, capacity planning platforms, alerting on occupancy drops and allocation problems.
  • Efficiency: measuring and improving how effectively every major workload uses the hardware it runs on, across training, inference, and evals.

Key Responsibilities:

  • Lead and grow the team, hire, onboard, coach, and retain senior and staff engineers.
  • Champion internal customers, engage with them directly, and bring what you learn back into the roadmap.
  • Own the roadmap, translate company-level compute strategy into a prioritized engineering roadmap.
  • Set the technical bar, review designs, weigh in on architecture, and hold the team to production standards.
  • Run the team as a product organization, ensure the team gathers its own requirements, defines schema contracts, and designs for a wide range of consumers.
  • Be the primary partner for cross-functional stakeholders, work closely with infrastructure, inference, research engineering, and finance leadership.
  • Drive operational excellence, own reliability and incident response for load-bearing systems.
  • Scale the function, anticipate where the team needs to grow in headcount, skills, and systems.

What You Bring:

  • Experience managing software or infrastructure engineering teams.
  • Strong technical background in production systems, data engineering, infrastructure, distributed systems, or observability.
  • Familiarity with at least one major cloud provider, Kubernetes-based infrastructure, and modern observability stacks.
  • Track record of setting and executing an engineering roadmap in an ambiguous, high-autonomy environment.
  • Excellent communication skills.
  • Comfort owning operational responsibility for systems the company depends on.

Preferred Qualifications:

  • Experience leading teams working on capacity planning, resource management, product engineering, or FinOps at a hyperscaler or in a large-scale ML environment.
  • Familiarity with accelerator infrastructure, GPU metrics, TPU utilization, or ML training and inference systems.
  • Experience with multi-cloud billing and telemetry normalization.
  • Experience building or leading internal data products with self-service access, schema contracts, and documentation.
  • Background in scheduling, packing efficiency, or profiling-driven optimization of large distributed workloads.

The annual compensation range for this role is $405,000-$485,000 USD.

This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/anthropic/jobs/5363210008