New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Anthropic

Research Product Manager, Model Behaviors

Anthropic
Apply →
hybrid senior full-time $385,000-$460,000 USD San Francisco, CA

First indexed 23 Jun 2026

Description

As a Product Manager for Model Behaviors at Anthropic, you will partner with the Alignment Finetuning team to define and shape Claude's character, behaviors, and reinforcement signals. This role directly influences how millions of people experience AI. You will systematically identify high-priority behavioral improvements, coordinate across Research, Product, and Safeguards teams, and accelerate the company's ability to ship well-aligned models.

Responsibilities:

  • Define behavioral defaults and steerability constraints
  • Develop and maintain taxonomies of model behaviors across capabilities
  • Identify, triage, and prioritize behavior issues and opportunities, coordinating input from Users, Research, Product, and Safeguards teams
  • Amplify alignment research breakthroughs, translating them into product, process, and model improvements
  • Deeply understand user interaction patterns to identify behavior improvements that make Claude more helpful and safe
  • Contribute to evals that measure alignment progress
  • Identify and scale initiatives and tools that help researchers ship alignment improvements faster

Minimum Qualifications:

  • Deep passion and curiosity for AI and LLMs, using AI regularly
  • 5+ years in product management leading scaled conversational AI products
  • First-principles thinker with the ability to navigate and execute amidst ambiguity
  • Track record of delivering products and features to end-users
  • Strong user empathy and ability to synthesize vague or contradictory feedback into actionable priorities
  • Strong judgment and model taste, with the ability to make tradeoffs when there is no clear right answer
  • Strong grasp of ML concepts and willingness to go deep on technical solutions
  • Intellectual curiosity without ego, comfortable asking questions and learning independently
  • Creative, hacker spirit and love solving puzzles

The annual compensation range for this role is $385,000-$460,000 USD.

This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/anthropic/jobs/5247407008