New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Anthropic

Safeguards Enforcement Analyst, User Well-being

Anthropic
Apply →
$245,000-$285,000 USD Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC

First indexed 4 Aug 2026

Description

As a Safeguards Analyst on the User Well-being team, you will focus on supporting the design and deployment of mental health guardrails – iterating on detection systems, managing review queues, evaluating new interventions, and monitoring existing ones.

The team covers a broad set of interconnected harms including suicide, self-harm, disordered eating, AI sycophancy, and emotional dependence on AI.

Key Responsibilities:

  • Support the design and execution of interventions, defining key metrics, and curating evaluation datasets
  • Partner with Engineering and Data Science teams to build, tune, and validate detection models for automated intervention systems
  • Monitor how interventions and detection systems perform over time
  • Review flagged content to drive enforcement and policy improvements
  • Support the development of in-product features that connect users to crisis resources
  • Support the Safeguards Policy Design team by providing detailed feedback on policy gaps
  • Keep up to date with emerging AI policy and external research on AI's relationship to mental health

Minimum Qualifications:

  • Experience in trust & safety, product policy, content moderation, or a related field
  • Experience designing or running experiments, evaluations, or measurement studies
  • Experience translating policy definitions into measurable form
  • Experience managing or coordinating content review operations
  • Proficiency in SQL and/or other data analysis tools
  • Experience working with generative AI products
  • Experience turning open questions and data into concise and insightful analysis
  • Understanding of the challenges involved in implementing product policies at scale
  • Sound judgment in ambiguous, high-consequence cases

Preferred Qualifications:

  • Subject matter expertise in mental health
  • Experience building or evaluating LLM-based classification systems
  • Experience using agentic tools to scale analysis or automate recurring work
  • Experience working within crisis support

The annual compensation range for this role is $245,000-$285,000 USD.

This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/anthropic/jobs/5374778008