# Engineering Manager, Infrastructure Engineering

**Company**: CoreWeave
**Location**: Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA
**Experience**: senior
**Job type**: full-time
**Salary**: $182,000 to $242,000
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/coreweave/jobs/4699712006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_b9c9869a-9bc

## Description

CoreWeave is seeking an experienced Engineering Manager to lead its MetalDev RAS team within Hardware Engineering Dev. The team focuses on the reliability, availability, and serviceability of CoreWeave's bare-metal infrastructure.

## Responsibilities

### People Leadership & Team Development

- Build, coach, and grow a team of infrastructure and site reliability engineers.

- Foster a culture of ownership and reliability.

- Run 1:1s, performance reviews, and career development.

### Reliability, Availability & Serviceability (RAS) Leadership

- Own the reliability, availability, and serviceability of MetalDev RedFish services.

- Define and drive KPIs, SLAs, and SLOs for the team.

- Champion system observability and health using tools like Prometheus and Grafana.

### Incident Management & Operational Excellence

- Establish and improve incident response processes.

- Lead communication to stakeholders during major incidents.

- Oversee the full server hardware lifecycle through automation and CI/CD pipelines.

### Cross Functional Collaboration

- Collaborate with engineering teams on platform reliability and resilience improvements.

- Represent the team in planning and prioritization.

- Engage with upstream communities and guide the team's technical standards.

## Requirements

- 3+ years of engineering management experience.

- 7+ years of combined experience in cloud operations, site reliability engineering, infrastructure, or related technical roles.

- Strong understanding of cloud platforms, K8S, and bare-metal infrastructure.

- Experience with incident management practices and frameworks.

- Experience leading teams that develop software in Go or comparable systems languages.

## Preferred Qualifications

- Experience managing bare-metal, hardware, or hardware-adjacent infrastructure teams.

- Familiarity with Redfish, BMC, or server lifecycle management technologies.

- Experience scaling teams, tooling, and processes in a high-growth environment.

## Compensation

The base salary range for this role is $182,000 to $242,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program.

## Skills

### Required
- cloud platforms
- K8S
- bare-metal infrastructure
- incident management
- Go
- Prometheus
- Grafana

### Nice to have
- Redfish
- BMC
- server lifecycle management technologies

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/coreweave/jobs/4699712006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
