New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
CoreWeave

Senior Systems Engineer, Test Frameworks & Validation Platform

CoreWeave
Apply →
senior full-time $153,000 to $204,000 Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA

First indexed 30 Jul 2026

Description

CoreWeave is The Essential Cloud for AI. As a Senior Systems Engineer on the Systems Engineering team, you will own and evolve our test framework, a Kubernetes-native system that qualifies host images and infrastructure on real hardware.

Responsibilities:

  • Own and extend the test framework that qualifies host images and infrastructure on real hardware
  • Broaden what the framework qualifies, including HPC and fabric verification, Slurm-on-Kubernetes, and further down the stack into firmware and hardware
  • Keep the pipeline fast, hermetic, and trustworthy, ensuring engineers trust the signal
  • Own the results and reporting path end to end, including structured test results, storage, dashboards, and triage surfaces
  • Explore AI-native testing, including LLM-driven log triage, failure classification, and regression detection
  • Collaborate with firmware, kernel, imaging, and HPC teams to embed testing into their release process

Requirements:

  • 3+ years of experience building test infrastructure, systems software, or platform tooling at scale
  • Fluent in Python, with either proven Rust experience or a strong systems background
  • Comfortable operating in a Kubernetes environment and reasoning about software deployment and testing
  • Solid Linux systems background, including boot chain, kernel and drivers, and low-level debugging
  • Real testing discipline, with strong opinions about flakiness, hermeticity, and signal
  • Clear communicator who treats the test framework as a product other engineers want to use

Preferred:

  • Rust and Kubernetes-native workflow orchestration experience
  • HPC or large-cluster experience, including InfiniBand/RoCE, GPU/accelerator validation, or performance-regression frameworks
  • Slurm or Slurm-on-Kubernetes experience
  • Firmware or lower-level hardware validation experience
  • Applying LLMs to test workflows, including triage, failure classification, and flaky-test detection
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/coreweave/jobs/4700154006