Description
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society.
The Domain Scaling team aims to make Claude world-class at real-world knowledge work in domains like finance, healthcare, and legal.
Responsibilities
- Own the data strategy for knowledge work verticals end-to-end, from task sourcing through RL training
- Manage technical relationships with external data vendors, including evaluation of data quality and reward design
- Collaborate with domain experts to design data pipelines and evaluations
- Explore novel ways of creating RL environments for high-value tasks
- Develop and improve QA frameworks to catch reward hacking and ensure environment quality
- Run generalization experiments to measure how data strategy changes improve model capabilities
- Partner with other RL research teams and product teams to translate capability goals into training environments and evaluations
Requirements
- Experience with fine-tuning large language models for specific domains or real-world use cases
- Experience with reinforcement learning, reward design, or training data curation for LLMs
- Comfortable managing technical vendor relationships and iterating quickly on feedback
- Strong cross-functional collaboration skills
- Passionate about making AI more useful and accessible across different industries
Nice to Have
- Experience training production ML systems
- Experience designing evaluations or benchmarks for LLMs
- Domain expertise in a vertical where Anthropic would like to make its models more useful
- Experience working with external vendors or technical partners
The annual compensation range for this role is $1-$2 USD.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/anthropic/jobs/5271380008