# Research Engineer, RL Scaling Science

**Company**: Anthropic
**Location**: London
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Salary**: £375,000-£640,000 GBP
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q116758847

**Apply**: https://job-boards.greenhouse.io/anthropic/jobs/5264619008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_c5b9c08f-47a

## Description

Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

As a Research Engineer on Anthropic's RL Scaling Science team, you will design and run large-scale experiments to understand and resolve bottlenecks, build benchmarks to measure long-horizon progress, and ship validated findings directly into production training.

This role sits at the boundary between research and engineering, with open problems, frontier-scale experiments, and a short path from results to production.

## Key Responsibilities

- Design, run, and interpret large-scale RL experiments, analysing data rigorously

- Investigate how RL improves with horizon, compute, and model size growth

- Build and maintain benchmarks for long-horizon RL to ensure measurable progress

- Translate validated findings into production training recipes

- Debug complex issues at the research-infrastructure seam

- Collaborate with adjacent RL teams across research and engineering

## Minimum Qualifications

- Strong empirical research skills in Reinforcement Learning or related areas

- Ability to own large experiments end-to-end

- Proficiency in Python and experience with large-scale ML systems

- Comfort operating at the research-systems boundary

- Concern for the societal impacts of AI and responsible scaling

## Preferred Qualifications

- Published or shipped work in long-horizon RL or RL fundamentals

- Experience translating research into production training recipes

- Demonstrated large-scale industry impact via RL interventions

- Experience with frontier-scale training runs

## Representative Projects

- Designing benchmark suites for long-horizon RL

- Stress-testing experimental findings across model scales

- Investigating unexpected scaling trends in RL runs

The annual compensation range for this role is £375,000-£640,000 GBP.

## Logistics

- Minimum education: Bachelor’s degree or equivalent

- Required field of study: Relevant to the role

- Location-based hybrid policy: 25% office time

- Visa sponsorship: Available

## Benefits

Anthropic offers competitive compensation, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space.

## Skills

### Required
- Reinforcement Learning
- Python
- large-scale ML systems
- empirical research

### Nice to have
- long-horizon RL
- production training recipes
- frontier-scale training runs

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/anthropic/jobs/5264619008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
