Description
About Anthropic ---------------- Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole.
About the Role --------------- The Inference System Dynamics team is responsible for understanding the whole system and holding it to a high bar across four dimensions: throughput, latency, reliability, and correctness. We measure how the fleet performs against its theoretical performance frontier, run cross-layer investigations to explain the gaps, and own the correctness checks that make sure Claude's outputs are right, not just fast, across hardware platforms and serving configurations.
Key Responsibilities -------------------
- Run cross-layer performance investigations across throughput, latency, and reliability, sizing the gap between actual fleet performance and theoretical rooflines, identifying root causes, and quantifying the value of closing them
- Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, numerics, and serving configurations, and lead the investigation when it catches a regression
- Build the observability, dashboards, and modeling tools that make throughput, latency, cost, reliability, correctness, and their interactions legible across the stack
- Partner with kernel, serving, routing, autoscaling, and capacity teams to prioritize and land the highest-impact optimizations your analysis surfaces
Minimum Qualifications ---------------------
- Hands-on performance engineering experience: profiling, roofline analysis, latency/throughput optimization, and root-cause investigation in complex production systems
- Proficiency in Python, with the ability to read, instrument, and contribute to large production codebases you didn’t write
- Solid data analysis skills (e.g. SQL, pandas, or similar) sufficient to turn raw telemetry into clear findings
- Ability to communicate quantitative results clearly in writing to influence priorities on teams you don't manage
Preferred Qualifications -----------------------
- Experience with ML systems, especially training or inference infrastructure or general LLL serving stacks. Direct large-scale inference experience is a strong plus
- Familiarity with GPU/TPU/accelerator performance concepts (memory bandwidth, kernel overheads, quantization, collective communication). Reasoning about these matters more than having written kernels yourself
- Experience with reliability engineering for high-throughput services: autoscaling, load balancing, request routing, tail latency
- Experience with model evaluation or numerical regression-detection pipelines
- Experience building observability or telemetry for distributed systems
Representative Projects ------------------------
- Trace a 350ms latency gap on a new accelerator platform from end-to-end request timing down to a server scheduling overhead, quantify the win, and land the fix directly or with the owning team
- Redesign the correctness eval gate: determine which signals reliably catch real model-output regressions versus noise, and make it the trusted release criterion across hardware backends
- Build a FLOPs funnel that breaks down where compute actually goes across the fleet, exposing the gap between achieved throughput and kernel rooflines
- Root-cause a numerical divergence between two hardware platforms to a specific kernel change, and define the acceptance threshold going forward
- Model the latency–cost impact of changing batch-sizing and utilization targets, and turn the result into the signal the autoscaler uses in production
Experience Level: senior Employment Type: full-time Workplace Type: hybrid Category: Engineering Industry: Technology Salary Range: $350,000-$850,000 USD Salary Min: 350000 Salary Max: 850000 Salary Currency: USD Salary Period: year Required Skills:
- Performance engineering
- Python
- Data analysis
- Communication
Preferred Skills:
- ML systems
- GPU/TPU/accelerator performance
- Reliability engineering
- Model evaluation
- Observability building