Description
As a Safeguards Analyst on the User Well-being team, you will focus on supporting the design and deployment of mental health guardrails – iterating on detection systems, managing review queues, evaluating new interventions, and monitoring existing ones.
The team covers a broad set of interconnected harms including suicide, self-harm, disordered eating, AI sycophancy, and emotional dependence on AI.
Key Responsibilities:
- Support the design and execution of interventions, defining key metrics, and curating evaluation datasets
- Partner with Engineering and Data Science teams to build, tune, and validate detection models for automated intervention systems
- Monitor how interventions and detection systems perform over time
- Review flagged content to drive enforcement and policy improvements
- Support the development of in-product features that connect users to crisis resources
- Support the Safeguards Policy Design team by providing detailed feedback on policy gaps
- Keep up to date with emerging AI policy and external research on AI's relationship to mental health
Minimum Qualifications:
- Experience in trust & safety, product policy, content moderation, or a related field
- Experience designing or running experiments, evaluations, or measurement studies
- Experience translating policy definitions into measurable form
- Experience managing or coordinating content review operations
- Proficiency in SQL and/or other data analysis tools
- Experience working with generative AI products
- Experience turning open questions and data into concise and insightful analysis
- Understanding of the challenges involved in implementing product policies at scale
- Sound judgment in ambiguous, high-consequence cases
Preferred Qualifications:
- Subject matter expertise in mental health
- Experience building or evaluating LLM-based classification systems
- Experience using agentic tools to scale analysis or automate recurring work
- Experience working within crisis support
The annual compensation range for this role is $245,000-$285,000 USD.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/anthropic/jobs/5374778008