# Senior Systems Engineer,  Test Frameworks & Validation Platform

**Company**: CoreWeave
**Location**: Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA
**Experience**: senior
**Job type**: full-time
**Salary**: $153,000 to $204,000
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/coreweave/jobs/4700154006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_66e88481-30b

## Description

CoreWeave is The Essential Cloud for AI. As a Senior Systems Engineer on the Systems Engineering team, you will own and evolve our test framework, a Kubernetes-native system that qualifies host images and infrastructure on real hardware.

Responsibilities:

- Own and extend the test framework that qualifies host images and infrastructure on real hardware

- Broaden what the framework qualifies, including HPC and fabric verification, Slurm-on-Kubernetes, and further down the stack into firmware and hardware

- Keep the pipeline fast, hermetic, and trustworthy, ensuring engineers trust the signal

- Own the results and reporting path end to end, including structured test results, storage, dashboards, and triage surfaces

- Explore AI-native testing, including LLM-driven log triage, failure classification, and regression detection

- Collaborate with firmware, kernel, imaging, and HPC teams to embed testing into their release process

Requirements:

- 3+ years of experience building test infrastructure, systems software, or platform tooling at scale

- Fluent in Python, with either proven Rust experience or a strong systems background

- Comfortable operating in a Kubernetes environment and reasoning about software deployment and testing

- Solid Linux systems background, including boot chain, kernel and drivers, and low-level debugging

- Real testing discipline, with strong opinions about flakiness, hermeticity, and signal

- Clear communicator who treats the test framework as a product other engineers want to use

Preferred:

- Rust and Kubernetes-native workflow orchestration experience

- HPC or large-cluster experience, including InfiniBand/RoCE, GPU/accelerator validation, or performance-regression frameworks

- Slurm or Slurm-on-Kubernetes experience

- Firmware or lower-level hardware validation experience

- Applying LLMs to test workflows, including triage, failure classification, and flaky-test detection

## Skills

### Required
- Python
- Rust
- Kubernetes
- Linux
- testing

### Nice to have
- HPC
- Slurm
- firmware
- LLM
- Argo Workflows

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/coreweave/jobs/4700154006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
