# Safeguards Enforcement Analyst, Violence & Extremism

**Company**: Anthropic
**Location**: Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Salary**: $285,000-$330,000 USD
**Category**: Operations
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q116758847

**Apply**: https://job-boards.greenhouse.io/anthropic/jobs/5343907008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_9cfe1b55-e67

## Description

As a Safeguards Enforcement Analyst focused on Violence & Extremism, you will build and execute operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas.

Your work spans detecting and mitigating attempts to misuse Anthropic's AI systems to facilitate real-world harm, including weapons and dangerous technology, critical infrastructure attacks, violent extremism, and threats of violence.

**Key Responsibilities:**

- Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy

- Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements

- Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations

- Review flagged content to drive enforcement decisions and surface policy gaps, with particular attention to novel or technically sophisticated misuse attempts and emerging extremist movements, ideologies, and mobilization tactics

- Support the Safeguards policy design team by providing structured feedback on policy gaps and enforcement ambiguities based on real enforcement scenarios

- Develop and maintain enforcement guidelines and reviewer documentation that enable accurate, consistent enforcement across a wide range of content

- Keep up to date with emerging threats, terrorist and extremist movements, regulatory changes, and AI policy enforcement best practices, and apply these to inform our workflows and evals

- Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent extremist activity

**Minimum Qualifications:**

- Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field, with direct exposure to harmful content, dangerous technology, violent extremism, or physical harm facilitation

- Experience standing up and scaling policy enforcement or content review workflows

- Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health

- Experience identifying emerging risks and threat actors, and communicating findings to a diverse set of stakeholders, such as Product, Policy, Engineering, and Legal teams

- Experience working with generative AI products, including writing effective prompts for content review and enforcement

- Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space

**Preferred Qualifications:**

- Subject matter expertise in one or more high-stakes harm areas, such as weapons and dangerous technology, violent extremism, terrorism, autonomous systems, or critical infrastructure protection

- Familiarity with relevant legal and regulatory frameworks governing dangerous technology, critical infrastructure, or domestic/international terrorism

- Experience developing evals or red-teaming AI systems, particularly for harmful content or policy enforcement use cases

- Experience with threat actor profiling and threat intelligence frameworks (e.g., MITRE ATT&CK)

- Experience tracking threat actors, extremist networks, or misuse patterns across surface, deep, and dark web environments

- Experience with large language models and an understanding of how AI technology could provide meaningful uplift toward serious harm

- Proficiency in Python for data analysis and workflow automation

- Background in law enforcement, national security, defense, counterterrorism, or a relevant regulatory environment

- Experience assessing the technical plausibility and real-world harm potential of content, including the ability to distinguish between general educational content and genuine operational uplift, and between protected speech and genuine incitement/mobilization

- Familiarity with cross-platform threat analysis and OSINT techniques

The annual compensation range for this role is $285,000-$330,000 USD.

## Skills

### Required
- SQL
- data analysis
- policy enforcement
- threat intelligence
- counterterrorism
- generative AI

### Nice to have
- Python
- MITRE ATT&CK
- large language models
- OSINT techniques

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/anthropic/jobs/5343907008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
