Description
Compensation
The base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. The salary range for this position is $250K – $445K, with generous equity, performance-related bonuses, and the following benefits:
- Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts
- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit)
- 401(k) retirement plan with employer match
- Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks)
- Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees
- 13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time
- Mental health and wellness support
- Employer-paid basic life and disability coverage
- Annual learning and development stipend to fuel your professional growth
- Daily meals in our offices, and meal delivery credits as eligible
- Relocation support for eligible employees
- Additional taxable fringe benefits, such as charitable donation matching and wellness stipends
About the Team
The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We focus on measuring monitorability, training mechanisms that affect monitorability, and methods to improve monitorability.
About the Role
We're looking for a researcher with strong empirical ML expertise and a deep interest in model behavior, alignment, or interpretability. As a researcher on the Alignment team, you will design and run experiments to improve our understanding of model monitorability, investigate how training interventions influence monitorability, and translate findings into practical oversight and training recommendations.
In this role, you will:
- Design and run empirical studies of chain-of-thought monitorability across frontier reasoning models and training settings
- Build evaluations that measure whether monitors can reliably predict properties of interest, including high-stakes forms of misbehavior
- Investigate how pre-training, synthetic data, mid-training, post-training, reinforcement learning, and other interventions improve or degrade monitorability
- Analyze model behavior and turn observations from monitoring into hypotheses, experiments, and recommendations
- Translate research findings into practical monitoring and oversight approaches that can inform real training runs
- Collaborate with researchers and engineers across model training, alignment evaluations, monitoring, and frontier-risk work
- Produce externally publishable research when results advance the broader science of alignment
You might thrive in this role if you:
- Have strong hands-on experience training, evaluating, or debugging large ML models, especially LLMs
- Have deep curiosity, interest in alignment, and high agency
- Bring depth in alignment, interpretability, model behavior, empirical ML, or adjacent research
- Are excited to investigate chain-of-thought monitorability, monitoring methods, and scalable oversight
- Can turn ambiguous research questions into measurable experiments and follow the evidence when results are subtle or noisy
- Move comfortably between research ideation and engineering execution
- Are curious about multiple approaches to understanding model behavior and are not committed to only one methodological lens
- Operate with high independence while collaborating closely across research and engineering teams
- Care about making increasingly capable AI systems more monitorable, trustworthy, and safe