# Staff Software Engineer, AI Reliability Engineering

**Company**: Anthropic
**Location**: London
**Work arrangement**: hybrid
**Experience**: staff
**Job type**: full-time
**Salary**: £325,000-£390,000 GBP
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q116758847

**Apply**: https://job-boards.greenhouse.io/anthropic/jobs/5101173008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_f2611d1b-e30

## Description

## Job Overview

We're seeking a Staff Software Engineer for our AI Reliability Engineering (AIRE) team at Anthropic, a pioneering AI organisation. You'll play a crucial role in ensuring the reliability of our AI systems, particularly Claude, which numerous users depend on.

## Responsibilities

- Develop Service Level Objectives for large language model serving systems, striking a balance between availability, latency, and development velocity.

- Design and implement monitoring and observability systems across the token path.

- Collaborate on designing and implementing high-availability serving infrastructure across multiple regions and cloud providers.

- Lead incident response for critical AI services, focusing on rapid recovery, thorough incident reviews, and systematic improvements.

- Support the reliability of safeguard model serving, essential for site reliability and Anthropic's safety commitments.

## Requirements and Fit

You should have:

- A strong background in distributed systems, infrastructure, or reliability.

- Experience as a reliability-minded software engineer or SRE.

- Curiosity and bravery, with comfort in navigating unfamiliar systems during incidents.

- Holistic thinking about system composition and identifying potential issues.

- Excellent communication and collaboration skills for effective teamwork across the organisation.

- Diverse experience in areas such as product stack development, database scaling, and distributed systems operation.

## Nice to Have

You may also have:

- Experience as an SRE, Production Engineer, or in similar reliability-focused roles.

- Hands-on experience operating large-scale model serving or training infrastructure (>1000 GPUs).

- Familiarity with ML hardware accelerators (GPUs, TPUs, Trainium).

- Knowledge of ML-specific networking optimisations like RDMA and InfiniBand.

- Expertise in AI-specific observability tools and frameworks.

- Experience with chaos engineering and systematic resilience testing.

- Contributions to open-source infrastructure or ML tooling.

## Logistics

- Annual Salary: £325,000-£390,000 GBP

- Minimum education: Bachelor’s degree or equivalent combination of education, training, and/or experience.

- Required field of study: Relevant to the role, demonstrated through coursework, training, or professional experience.

- Location-based hybrid policy: Office attendance expected at least 25% of the time.

- Visa sponsorship: Available, with efforts made to secure visas for successful candidates.

## Benefits

Anthropic offers competitive compensation and benefits, including optional equity donation matching, generous vacation and parental leave, flexible working hours, and a welcoming office environment.

## Culture

At Anthropic, we value big science and collaborative research. We work on a few large-scale research efforts, prioritising impact and advancing our long-term goals of steerable, trustworthy AI. Effective communication skills are highly valued.

## Skills

### Required
- distributed systems
- infrastructure
- reliability engineering
- large language models
- monitoring and observability
- high-availability serving infrastructure
- incident response
- ML hardware accelerators
- AI-specific observability tools

### Nice to have
- SRE
- Production Engineering
- chaos engineering
- systematic resilience testing
- open-source infrastructure
- ML tooling
- RDMA
- InfiniBand

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/anthropic/jobs/5101173008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
