Description
Anthropic's mission is to create reliable, interpretable, and steerable AI systems. You will be a foundational member of the Product Prompt and Eval Design team, focusing on building evals that check prompts, the harness that runs them, and tools for designers.
Key responsibilities:
- Write and revise prompts behind Claude's tools, features, and behaviours
- Build graders that prove a prompt fix and rerun on the next model
- Develop visual, low-code eval tools for designers
- Support model releases and maintain the eval harness
- Package what prompting can't fix for training
Minimum qualifications:
- Production-quality Python
- Experience with evaluation pipelines for LLM products
- Experience building internal tools for non-coders
- Experience with test harnesses and prompt engineering
Preferred qualifications:
- Experience within a model-launch cycle
- A/B testing experience
- Front-end or notebook-to-app experience
- Turning product rubrics into training signals
The annual compensation range for this role is $305,000-$385,000 USD.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/anthropic/jobs/5411318008