# Applied AI Engineer, Inference

**Company**: CoreWeave
**Location**: Bellevue, WA
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Salary**: $188,000 to $275,000
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/coreweave/jobs/4663228006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_24fc4923-759

## Description

CoreWeave is The Essential Cloud for AI. We're looking for an Applied AI Engineer to help us understand, measure, and improve the real-world performance of our inference platform.

The Inference team is responsible for delivering high-performance model serving capabilities that meet the needs of real production workloads. The Applied AI Engineer will focus on building and running rigorous benchmarks, profiling model and system behavior, identifying bottlenecks, and driving targeted optimizations.

Responsibilities:

- Build and maintain benchmarking workflows that measure latency, throughput, quality regressions, and cost across priority models and serving configurations.

- Benchmark the inference stack against realistic customer workloads and external provider baselines.

- Profile model-serving behavior across frameworks, runtimes, and hardware configurations.

- Drive targeted optimization efforts for specific customer and product workloads.

- Design and run experiments on model-serving techniques such as quantization and speculative decoding.

- Partner with inference platform engineers to productionize improvements.

- Produce clear technical writeups and recommendations.

Requirements:

- 4+ years of experience in machine learning, systems, performance engineering, or adjacent applied engineering work.

- Strong programming skills in Python.

- Experience running empirical evaluations, benchmarks, or experiments.

- Familiarity with LLM inference systems and tools.

- Understanding of practical tradeoffs involved in latency, throughput, batching, GPU utilization, quantization, and quality regression analysis.

Preferred:

- Experience optimizing inference workloads on modern GPU hardware.

- Experience with profiling tools.

- Familiarity with benchmark suites and evaluation frameworks.

- Experience using real production traces or customer traffic patterns.

We offer a competitive salary, equity awards, and a comprehensive benefits program.

## Skills

### Required
- Python
- machine learning
- systems
- performance engineering
- LLM inference systems

### Nice to have
- GPU hardware
- profiling tools
- benchmark suites
- production traces

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/coreweave/jobs/4663228006?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
