New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Anthropic

AI Infrastructure Operations, Demand Planning

Anthropic
Apply →
hybrid senior full-time $320,000-$405,000 USD San Francisco, CA

First indexed 11 Aug 2026

Description

Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

The AI Infrastructure Operations, Demand Planning role sits in the Planning pillar, on the Demand Planning team, and works daily with research engineering, pretraining, inference, compute supply, finance, and external vendors.

Responsibilities:

  • Turn the forecast into per-tranche requirements. Take the Demand Planning forecast plus direct input from research, pretraining, and inference planners, and convert it into concrete accelerator, interconnect, region, supporting-resource, and date requirements for each tranche. Represent those in sourcing negotiations and data center build reviews, including which contractual terms actually move delivery dates.
  • Qualify tranches for deliverability before signature. The Capacity Planner signs fit-to-forecast; you sign whether the shape can land schedulable, healthy, and instrumented in that region on that date, with storage, egress, identity in place.
  • Close the delivery loop. Track forecast-versus-delivered on shape, region, and timing for every tranche; publish the variance; and feed it back to Demand Planning and into the next contract.
  • Own the bring-up system of record. Define the canonical contract-to-occupied state machine with explicit entry and exit criteria per stage, and make it a first-class object in the capacity data layer so every downstream tool sees in-flight capacity, not only what has landed.
  • Run a portfolio of bring-ups in parallel , new cloud regions, on-prem sites, neocloud blocks , with one integrated schedule spanning provider milestones, cluster creation, network turn-up, storage readiness, health burn-in, and first-workload landing.
  • Drive readiness automation: All capacity systems are fully integrated for all new capacity, from contracted through ingested, automated and scaled.
  • Instrument and publish the numbers that matter , time-to-occupied and paid-idle dollars per tranche , with executive-level reporting on status, tradeoffs, and risk across the portfolio.

Requirements:

  • Significant experience delivering large-scale infrastructure , cloud regions, accelerator clusters, HPC systems, or bare-metal fleets , at multi-region scale or ≥10k accelerators (or CPU/storage equivalent).
  • Technical range from through cluster orchestration and node health, up to the telemetry and planning tables on top , enough to debug where they disagree rather than route it.
  • SQL and enough Python to answer your own questions and build your own reporting.
  • A degree in a technical field or an equivalent engineering track record.

Preferred Skills:

  • Reserved-capacity onboarding, private offers, or capacity commitments with cloud or neocloud providers.
  • Enough demand-planning exposure to challenge a forecast, translate it into per-tranche requirements, and feed delivery variance back into it.
  • Data center or colocation delivery: power and space planning, network turn-up, site acceptance, vendor management.
  • Accelerator health and burn-in, collective-communications sanity testing, or fleet-health SLOs , and a rigorous definition of "healthy."
  • Systems of record or lifecycle services for infrastructure assets.
  • Onboarding a new hardware generation into an existing scheduler and observability stack.

The annual compensation range for this role is $320,000-$405,000 USD.

This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/anthropic/jobs/5382750008