# Performance Engineer, Inference Systems

**Company**: Anthropic
**Location**: San Francisco, CA | New York City, NY | Seattle, WA
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Salary**: $350,000-$850,000 USD
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q116758847

**Apply**: https://job-boards.greenhouse.io/anthropic/jobs/5224564008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_1c0486c2-3f5

## Description

About Anthropic ---------------- Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole.

About the Role --------------- The Inference System Dynamics team is responsible for understanding the whole system and holding it to a high bar across four dimensions: throughput, latency, reliability, and correctness. We measure how the fleet performs against its theoretical performance frontier, run cross-layer investigations to explain the gaps, and own the correctness checks that make sure Claude's outputs are right, not just fast, across hardware platforms and serving configurations.

Key Responsibilities -------------------

- Run cross-layer performance investigations across throughput, latency, and reliability, sizing the gap between actual fleet performance and theoretical rooflines, identifying root causes, and quantifying the value of closing them

- Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, numerics, and serving configurations, and lead the investigation when it catches a regression

- Build the observability, dashboards, and modeling tools that make throughput, latency, cost, reliability, correctness, and their interactions legible across the stack

- Partner with kernel, serving, routing, autoscaling, and capacity teams to prioritize and land the highest-impact optimizations your analysis surfaces

Minimum Qualifications ---------------------

- Hands-on performance engineering experience: profiling, roofline analysis, latency/throughput optimization, and root-cause investigation in complex production systems

- Proficiency in Python, with the ability to read, instrument, and contribute to large production codebases you didn’t write

- Solid data analysis skills (e.g. SQL, pandas, or similar) sufficient to turn raw telemetry into clear findings

- Ability to communicate quantitative results clearly in writing to influence priorities on teams you don't manage

Preferred Qualifications -----------------------

- Experience with ML systems, especially training or inference infrastructure or general LLL serving stacks. Direct large-scale inference experience is a strong plus

- Familiarity with GPU/TPU/accelerator performance concepts (memory bandwidth, kernel overheads, quantization, collective communication). Reasoning about these matters more than having written kernels yourself

- Experience with reliability engineering for high-throughput services: autoscaling, load balancing, request routing, tail latency

- Experience with model evaluation or numerical regression-detection pipelines

- Experience building observability or telemetry for distributed systems

Representative Projects ------------------------

- Trace a 350ms latency gap on a new accelerator platform from end-to-end request timing down to a server scheduling overhead, quantify the win, and land the fix directly or with the owning team

- Redesign the correctness eval gate: determine which signals reliably catch real model-output regressions versus noise, and make it the trusted release criterion across hardware backends

- Build a FLOPs funnel that breaks down where compute actually goes across the fleet, exposing the gap between achieved throughput and kernel rooflines

- Root-cause a numerical divergence between two hardware platforms to a specific kernel change, and define the acceptance threshold going forward

- Model the latency–cost impact of changing batch-sizing and utilization targets, and turn the result into the signal the autoscaler uses in production

Experience Level: senior Employment Type: full-time Workplace Type: hybrid Category: Engineering Industry: Technology Salary Range: $350,000-$850,000 USD Salary Min: 350000 Salary Max: 850000 Salary Currency: USD Salary Period: year Required Skills:

- Performance engineering

- Python

- Data analysis

- Communication

Preferred Skills:

- ML systems

- GPU/TPU/accelerator performance

- Reliability engineering

- Model evaluation

- Observability building

## Skills

### Required
- performance engineering
- Python
- data analysis
- communication

### Nice to have
- ML systems
- GPU/TPU/accelerator performance
- reliability engineering
- model evaluation
- observability building

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/anthropic/jobs/5224564008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
