# Agent Post-Training, Artifacts Research

**Company**: OpenAI
**Location**: San Francisco
**Experience**: senior
**Job type**: Full time
**Salary**: $295K - $445K
**Category**: Research
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q124605186

**Apply**: https://jobs.ashbyhq.com/openai/6897d024-88c1-43ed-adb8-5d2fc5eec984?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_5eb05828-829

## Description

We are seeking an Agent Post-Training, Artifacts Research to join our Research team in San Francisco.

The base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. The estimated base salary for this role is $295K – $445K.

## About the Team

The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve.

## About the Role

As a member of Agent Post-Training, Artifacts, you will train frontier models to create polished, useful work products: documents, spreadsheets, slide decks, dashboards, reports, analyses, and other interactive or editable artifacts. You will help teach our models to move from a vague user goal to a finished artifact with strong structure, visual taste, domain judgment, correctness, and low latency.

## Responsibilities

- Design and run experiments that improve agentic model behavior for complex software and plugins.

- Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis.

- Build evals and environments that expose the next set of model failures, then turn those failures into training data, product fixes, or new research directions.

- Partner with Codex and ChatGPT product teams to understand what users need and translate product signal into model improvements.

- Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops that shape downstream agent behavior.

- Help decide which integrations, capabilities, and fixes are ready for inclusion in major model runs.

- Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.

- Take on cross-functional projects that touch model training, product infrastructure, and the production agent harness, such as multi-agent systems or training directly against production-like environments.

- Debug hard failures in shipped or near-shipped models and turn messy qualitative behavior into concrete hypotheses, experiments, and fixes.

## Benefits

- Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts.

- Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit).

- 401(k) retirement plan with employer match.

- Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks).

- Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees.

- 13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time.

- Mental health and wellness support.

- Employer-paid basic life and disability coverage.

- Annual learning and development stipend to fuel your professional growth.

- Daily meals in our offices, and meal delivery credits as eligible.

- Relocation support for eligible employees.

## Requirements

- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field.

- Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.

- Experience with designing and running experiments, and analyzing results.

- Ability to work across research, product, infrastructure, data, evals, and safety boundaries.

## Skills

### Required
- machine learning
- software engineering
- systems
- statistics
- LLMs
- RL
- RLHF/RLAIF
- post-training
- evals
- graders
- synthetic data
- model training

---

Source: [Apply at jobs.ashbyhq.com](https://jobs.ashbyhq.com/openai/6897d024-88c1-43ed-adb8-5d2fc5eec984?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
