Description
You will build and run machine learning experiments to understand and steer powerful AI systems, focusing on safety and beneficial outcomes.
As a Research Engineer on Alignment Science, you'll contribute to exploratory experimental research on AI safety, particularly risks from future powerful systems.
Responsibilities include:
- Designing and running experiments to test AI safety techniques
- Collaborating with teams like Interpretability, Fine-Tuning, and Frontier Red Team
- Contributing to research papers, blog posts, and talks
- Building tooling to evaluate novel LLM-generated jailbreaks
- Writing scripts and prompts to test models' reasoning abilities
Requirements:
- Significant software, ML, or research engineering experience
- Experience contributing to empirical AI research projects
- Familiarity with technical AI safety research
- Ability to work collaboratively and pick up slack
- Care about AI impacts
Strong candidates may also have:
- Experience authoring research papers in machine learning, NLP, or AI safety
- Experience with LLMs, reinforcement learning, or Kubernetes clusters
The annual compensation range for this role is £260,000-£370,000 GBP.
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/anthropic/jobs/4610158008