# Engineering Manager, Kubernetes Infrastructure (Bare Metal)

**Company**: CoreWeave
**Location**: Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA
**Experience**: senior
**Job type**: full-time
**Salary**: $182,000 to $242,000
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/coreweave/jobs/4703779006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_e794ce91-733

## Description

CoreWeave is seeking an Engineering Manager to lead a team building and operating Kubernetes infrastructure on bare metal. The team is responsible for the reliability, scalability, and operational excellence of systems powering high-performance AI and ML workloads.

You will lead engineers working on cluster lifecycle, platform reliability, infrastructure automation, and operational systems. This is a hands-on leadership role for someone who can grow engineers, improve execution, and partner with platform, networking, compute, and product teams.

Key responsibilities:

- Lead a team of engineers responsible for Kubernetes infrastructure running on bare metal

- Set clear goals, priorities, and execution plans for the team

- Partner with senior ICs and adjacent teams on the roadmap for cluster lifecycle management, upgrades, reliability, observability, and infrastructure automation

- Improve the team's operational excellence across incident response, on-call health, root-cause analysis, and service ownership

- Drive engineering best practices for safe change management, testing, rollout quality, and production readiness

Requirements:

- Experience managing an infrastructure, platform, or SRE-oriented engineering team

- Strong technical depth in Kubernetes, distributed systems, and production infrastructure

- Experience operating Kubernetes in complex environments, ideally including bare metal, hybrid, or highly performance-sensitive systems

Preferred Experience:

- Experience with GPU-heavy, HPC, or ML infrastructure environments

- Experience with bare-metal infrastructure, server lifecycle operations, or low-level systems troubleshooting

What success looks like:

- In the first 90 days: Build trust with the team and key partner organizations, assess team health, and establish core operating rhythms

- In the first 6 months: Improve predictability of team execution and service ownership, raise the quality bar for change management and rollout safety

- In the first 12 months: Build a strong, durable team with clear ownership and healthy operating mechanisms, deliver meaningful infrastructure improvements

The base salary range for this role is $182,000 to $242,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location.

## Skills

### Required
- Kubernetes
- distributed systems
- production infrastructure
- cluster lifecycle management
- infrastructure automation

### Nice to have
- GPU-heavy infrastructure
- HPC
- ML infrastructure
- bare-metal infrastructure
- server lifecycle operations

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/coreweave/jobs/4703779006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
