# Senior Technical Program Manager – Fleet Operations & Reliability

**Company**: CoreWeave
**Location**: New York, NY
**Experience**: senior
**Job type**: full-time
**Salary**: $157,000 to $210,000
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/coreweave/jobs/4710650006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_1c45e9a5-fba

## Description

CoreWeave is building the world's largest AI Cloud platform, and our fleet is growing at extraordinary speed. We're seeking a Senior Technical Program Manager to own fleet-wide reliability and operations programs that keep pace with that growth.

This role partners closely with Compute, Networking, Data Center, and Operations teams to drive improvements in fleet operating efficiency , maximizing the percentage of fleet capacity that is healthy, available, and sellable , alongside reliability and stability goals.

You'll act as the central owner for fleet reliability and operations programs: aligning teams on priorities, defining success metrics, driving execution, and ensuring improvements hold at scale , not just land once and regress as the fleet scales.

The Technical Program Manager will:

- Establish and own fleet reliability metrics and dashboards (e.g., failure rates, MTTR, incident trends, capacity availability and utilization rates), with visibility into how these trend as the fleet expands

- Drive alignment on fleet reliability OKRs across engineering and operations teams

- Identify systemic reliability gaps across hardware, firmware, software, networking, storage, and operational processes

- Lead complex, cross-functional programs to improve fleet delivery, readiness, and operational stability that scale with fleet growth, avoiding solutions that only work at current size

- Define program plans, milestones, dependencies, risks, and success criteria for reliability initiatives

- Proactively manage cross-team dependencies and unblock execution across multiple engineering organizations

- Track progress against goals, surface risks early, and communicate status clearly to stakeholders and leadership

- Participate in major incident reviews and root cause analysis, ensuring follow-up actions are tracked to closure

- Use data and post-incident learnings to prioritize reliability investments and drive corrective action

Minimum Qualifications:

- Bachelor's degree in Computer Engineering, Computer Science, or a related technical field, or equivalent practical experience

- 7+ years of technical program management experience in large-scale compute infrastructure, cloud, or platform environments

- Experience operating in a rapidly growing or fast-scaling infrastructure environment, with programs designed to hold up under 2x, 5x, or greater fleet growth

- Strong technical aptitude across infrastructure domains (compute, storage, networking, hardware, or SRE)

- Demonstrated ability to use data and metrics to drive prioritization, execution, and decision-making

- Excellent communication and stakeholder management skills, including executive-level reporting

Preferred Qualifications:

- Experience operating at scale in data center, cloud infrastructure, or hyperscale environments balancing reliability investments against capacity goals, with an understanding of how operational decisions affect sellable capacity

- Familiarity with reliability frameworks such as SLIs/SLOs, error budgets, incident management practices, and root cause analysis

- Understanding of hardware failure and processes, including failure diagnosis, vendor coordination, and replacement lifecycle management

- Background in observability, monitoring, or telemetry systems (e.g., Prometheus, Grafana, OpenTelemetry)

- Experience with fleet lifecycle management (provisioning, firmware/OS updates, decommissioning) at scale

The base salary range for this role is $157,000 to $210,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location.

Benefits:

- Medical, dental, and vision insurance - 100% paid for by CoreWeave

- Company-paid Life Insurance

- Voluntary supplemental life insurance

- Short and long-term disability insurance

- Flexible Spending Account

- Health Savings Account

- Tuition Reimbursement

- Ability to Participate in Employee Stock Purchase Program (ESPP)

- Mental Wellness Benefits through Spring Health

- Family-Forming support provided by Carrot

- Paid Parental Leave

- Flexible, full-service childcare support with Kinside

- 401(k) with a generous employer match

- Flexible PTO

- Catered lunch each day in our office and data center locations

- A casual work environment

- A work culture focused on innovative disruption

## Skills

### Required
- technical program management
- large-scale compute infrastructure
- cloud
- platform environments
- data analysis
- communication
- stakeholder management

### Nice to have
- reliability frameworks
- SLIs/SLOs
- error budgets
- incident management practices
- root cause analysis
- observability
- monitoring
- telemetry systems
- fleet lifecycle management

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/coreweave/jobs/4710650006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
