New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Anthropic

Performance Engineer, Inference Systems

Anthropic
Apply →
hybrid senior full-time $350,000-$850,000 USD San Francisco, CA | New York City, NY | Seattle, WA

First indexed 21 May 2026

Description

About Anthropic ---------------- Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole.

About the Role --------------- The Inference System Dynamics team is responsible for understanding the whole system and holding it to a high bar across four dimensions: throughput, latency, reliability, and correctness. We measure how the fleet performs against its theoretical performance frontier, run cross-layer investigations to explain the gaps, and own the correctness checks that make sure Claude's outputs are right, not just fast, across hardware platforms and serving configurations.

Key Responsibilities -------------------

  • Run cross-layer performance investigations across throughput, latency, and reliability, sizing the gap between actual fleet performance and theoretical rooflines, identifying root causes, and quantifying the value of closing them
  • Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, numerics, and serving configurations, and lead the investigation when it catches a regression
  • Build the observability, dashboards, and modeling tools that make throughput, latency, cost, reliability, correctness, and their interactions legible across the stack
  • Partner with kernel, serving, routing, autoscaling, and capacity teams to prioritize and land the highest-impact optimizations your analysis surfaces

Minimum Qualifications ---------------------

  • Hands-on performance engineering experience: profiling, roofline analysis, latency/throughput optimization, and root-cause investigation in complex production systems
  • Proficiency in Python, with the ability to read, instrument, and contribute to large production codebases you didn’t write
  • Solid data analysis skills (e.g. SQL, pandas, or similar) sufficient to turn raw telemetry into clear findings
  • Ability to communicate quantitative results clearly in writing to influence priorities on teams you don't manage

Preferred Qualifications -----------------------

  • Experience with ML systems, especially training or inference infrastructure or general LLL serving stacks. Direct large-scale inference experience is a strong plus
  • Familiarity with GPU/TPU/accelerator performance concepts (memory bandwidth, kernel overheads, quantization, collective communication). Reasoning about these matters more than having written kernels yourself
  • Experience with reliability engineering for high-throughput services: autoscaling, load balancing, request routing, tail latency
  • Experience with model evaluation or numerical regression-detection pipelines
  • Experience building observability or telemetry for distributed systems

Representative Projects ------------------------

  • Trace a 350ms latency gap on a new accelerator platform from end-to-end request timing down to a server scheduling overhead, quantify the win, and land the fix directly or with the owning team
  • Redesign the correctness eval gate: determine which signals reliably catch real model-output regressions versus noise, and make it the trusted release criterion across hardware backends
  • Build a FLOPs funnel that breaks down where compute actually goes across the fleet, exposing the gap between achieved throughput and kernel rooflines
  • Root-cause a numerical divergence between two hardware platforms to a specific kernel change, and define the acceptance threshold going forward
  • Model the latency–cost impact of changing batch-sizing and utilization targets, and turn the result into the signal the autoscaler uses in production

Experience Level: senior Employment Type: full-time Workplace Type: hybrid Category: Engineering Industry: Technology Salary Range: $350,000-$850,000 USD Salary Min: 350000 Salary Max: 850000 Salary Currency: USD Salary Period: year Required Skills:

  • Performance engineering
  • Python
  • Data analysis
  • Communication

Preferred Skills:

  • ML systems
  • GPU/TPU/accelerator performance
  • Reliability engineering
  • Model evaluation
  • Observability building
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/anthropic/jobs/5224564008