# Research Engineer, Takeoff Intel

**Company**: Anthropic
**Location**: San Francisco, CA
**Job type**: full-time
**Salary**: $350,000-$850,000 USD
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q116758847

**Apply**: https://job-boards.greenhouse.io/anthropic/jobs/5416882008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_439e9bf6-c5b

## Description

## About the Role

As a Research Engineer on Takeoff Intel you'll build and run the evaluation and measurement instruments that make this research possible. This is a generalist role on a small team: you'll work across evals infrastructure, large-scale data processing, and analysis tooling, and you'll prioritize shipping. We build instruments that answer real questions and help set priorities, not dashboards that surface noise. We value working prototypes, rapid iteration, accuracy and good prioritization. We often need to go from a vague research question to a running instrument quickly.

## Responsibilities

- Design, build, and run capability evaluations and measurement instruments at scale

- Build the data and analysis pipelines that turn large volumes of model outputs and telemetry into reliable metrics

- Prototype new instruments fast, validate them, and decide what to keep

- Review and supervise AI-written code as a normal part of the workflow

- Work closely with research scientists on the team and with partner teams to define what's worth measuring

- Contribute to internal write-ups and public reporting

## Requirements

- Have shipped an evaluation, data product, or research library end to end

- Prototype fast and are comfortable throwing code away

- Handle messy, large-volume data without over-engineering

- Have run experiments on large language models, not just moved their outputs around

- Can work from a vague question rather than a spec

- Communicate results clearly and collaborate closely with the researchers whose questions your instruments answer

## Nice to Have

- Built evaluation harnesses or benchmark infrastructure for LLMs

- Experience with large-scale ML or data infrastructure (self-driving, observability, or similar) alongside ML exposure

- Built tools or libraries that other researchers rely on

- A track record of catching what AI-written code gets wrong

## Logistics

The annual compensation range for this role is $350,000-$850,000 USD.

## Skills

### Required
- evaluation
- measurement instruments
- large-scale data processing
- analysis tooling
- AI-written code review

### Nice to have
- evaluation harnesses
- benchmark infrastructure for LLMs
- large-scale ML or data infrastructure
- tools or libraries for researchers

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/anthropic/jobs/5416882008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
