# Performance Engineer, Inference Engine

**Company**: Anthropic
**Location**: San Francisco, CA
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Salary**: $350,000-$850,000 USD
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q116758847

**Apply**: https://job-boards.greenhouse.io/anthropic/jobs/5418323008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_9cd7d6cf-5af

## Description

## Job Overview

We are seeking a Performance Engineer to work on our Inference Engine at Anthropic. The Inference Engine manages the token path between accelerator kernels and the routing layer, handling batching requests, model layout across chips, memory management, and coordination of forward passes.

## Responsibilities

You will work on building and optimizing the Inference Engine at Anthropic scale, focusing on improving throughput, cost, reliability, and latency across accelerator and cloud platforms. Key responsibilities include:

- Optimizing device utilization to prevent idle time due to overheads

- Implementing efficient reuse of model state to minimize recomputation

- Developing observability tools to identify bottlenecks, model improvements, and measure impact

- Ensuring model quality and efficiency without compromising robustness

- Collaborating with safety teams to integrate efficiency and safety features

## Requirements

- Strong understanding of LLM inference and its impact on accelerator compute, memory, and interconnect

- Proven ability to quickly learn complex systems and implement significant changes

- Expertise in systems programming (Rust, C++, or similar) with attention to code quality and testing

- Analytical approach to performance optimization

- Experience with pair programming and consideration for societal impacts

## Preferred Qualifications

- Experience with LLM serving engines and their abstractions

- GPU/accelerator programming expertise

- Knowledge of OS internals and language modeling with transformers

- Experience building allocators, caches, schedulers, or high-bandwidth transport

- Fluency in Rust and experience with systems reproducibility

## Logistics

- Annual salary: $350,000 - $850,000 USD

- Minimum education: Bachelor’s degree or equivalent

- Required field of study: Relevant to the role

- Location-based hybrid policy: 25% office time

- Visa sponsorship: Available

## Benefits

- Competitive compensation and benefits

- Optional equity donation matching

- Generous vacation and parental leave

- Flexible working hours

- Lovely office space

## About Anthropic

Anthropic is a public benefit corporation focused on creating reliable, interpretable, and steerable AI systems.

## Skills

### Required
- LLM inference
- systems programming
- performance optimization
- pair programming

### Nice to have
- LLM serving engines
- GPU/accelerator programming
- OS internals
- language modeling with transformers
- Rust

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/anthropic/jobs/5418323008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
