Description
Anthropic's Data Center Operations (DCO) team ensures compute fleet availability through hardware and IT operations. As the site lead, you own site outcomes for your assigned sites including: deployment velocity, availability, and incident response.
You will define the operational processes, quality gates, and governance rhythms for partner-operated sites. You provide tactical direction, set priorities, and define the standards for the vendor's on-site teams, paired with performance oversight and ongoing operational assessment to ensure all operational commitments are met.
Responsibilities:
- Own site availability, deployment milestones, and repair turnaround, verified with independent data rather than vendor self-reporting.
- Set daily and weekly priorities and lead the operating cadence, including standups and business reviews.
- Author and improve procedures for deployment, break-fix, change management, security, and EHS compliance. Analyze operational trends and standardize lessons across the program.
- Track vendor performance against SLAs and staffing commitments, driving corrective actions when necessary.
- Participate in the incident escalation on-call rotation. When designated Anthropic Incident Commander for a site-specific incident, direct vendor response, own communications, and close out post-incident actions.
- Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.
Requirements:
- 8+ years of experience in data center operations (hardware, IT infrastructure, or critical facilities) as a manager, technical lead or related role, including accountability for production availability.
- Managed vendors, MSPs, or contract workforces to measurable outcomes: SOWs, SLAs, operational reviews, and corrective action.
- Hands-on technical depth in server, network, and rack-level infrastructure, enough to independently verify vendor claims and audit quality.
- Built or substantially improved operational processes, not just run them.
- Served in an incident command or lead-responder role and communicate clearly under ambiguity.
- Can support non-standard hours, including an on-call rotation and availability during deployment surges and maintenance windows.
Benefits:
- Competitive compensation and benefits
- Optional equity donation matching
- Generous vacation and parental leave
- Flexible working hours
- Lovely office space
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/anthropic/jobs/5398280008