Description
Job Overview
As a Repairs Lead within the Data Center Infrastructure organisation, you will define and manage the end-to-end hardware repair program across Anthropic's growing fleet of data centres.
Responsibilities Include:
- Define the global repair strategy including repair SLAs, prioritisation rules, escalation paths, and reporting methods, and drive standardisation across all sites.
- Own repair turnaround time and repair backlog across the fleet evidenced by Anthropic-owned ticket and telemetry dashboards you help develop.
- Author and improve procedures for triage, break-fix, return-to-service validation, and train site operations partners on how to execute them.
- Manage RMA and reverse logistics programs with OEMs, ODMs, and depot repair vendors, including warranty claims, return cycle times, and failure analysis feedback.
- Set spares pool sizing and stocking levels by site and part, in coordination with supply chain and asset management, so that parts availability never gates repair SLAs.
- Analyse failure patterns across sites to identify root causes and drive corrective actions with hardware engineering, suppliers, and site operations owners.
- Lead the operating cadence with vendor and site leads, including weekly repair reviews, scorecards, and business reviews, and drive corrective actions for SLA excursions.
- Communicate repair constraints, risks, and fleet availability impact to engineering and leadership.
Requirements:
- Have 8+ years of experience in data centre operations as a manager, technical lead or related role, including accountability for production availability.
- Can demonstrate a proven track record running break-fix programs at a large scale across multiple sites.
- Have managed vendors, OEMs, or contract workforces to measurable outcomes: SLAs, operational reviews, and corrective action.
- Possess hands-on technical depth in server, network, and rack-level hardware, enough to independently verify repair quality and audit vendor claims.
- Have built or substantially improved operational processes.
- Are comfortable working with ticket, telemetry, and inventory data to drive decisions.
- Possess a bachelor's degree in relevant domain or equivalent practical experience.
Nice to Have:
- Experience with GPU/accelerator or high-density liquid-cooled infrastructure, including tray, cold plate, and manifold-level repair.
- Experience managing RMA and warranty programs with hyperscale OEMs/ODMs, including failure analysis and supplier quality engagement.
- Experience with spares planning, reverse logistics, or depot repair at data centre scale.
- Experience delivering repair outcomes inside partner-operated or colocation sites where on-the-floor operations are staffed via third parties.
- Familiarity with optics and high-speed interconnect failures.
Compensation:
The annual compensation range for this role is $320,000-$405,000 USD.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/anthropic/jobs/5399160008