# Staff Software Engineer, RL Environments

**Company**: Scale AI
**Location**: San Francisco, CA
**Experience**: staff
**Job type**: full-time
**Salary**: $252,000-$315,000 USD
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q112629176

**Apply**: https://job-boards.greenhouse.io/scaleai/jobs/4729820005?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_22c45f30-969

## Description

## Job Overview

As a Staff Software Engineer, RL Environments, you'll play a crucial role in building, running, verifying, and delivering RL environments at scale. You'll work on designing the platform for RL environments, including sandboxed execution, environment packaging, and rollout orchestration.

## Key Responsibilities

- Design and build the technical foundation for RL environments at scale

- Develop and maintain containerized worlds with real dependencies, state, tools, and graders

- Work on platform features such as sandboxed execution, environment packaging, and versioning

- Collaborate with engineers and domain experts to produce environments without reinventing infrastructure

- Instrument real applications, design task suites, and build graders that hold up under adversarial optimization

## Requirements

- 8+ years of software engineering experience with strong fundamentals in distributed systems, system design, data structures, and algorithms

- Strong Python skills and experience with shipping production software

- Comfort in at least one other part of the stack (TypeScript/React, Go, Rust, or similar)

- Deep experience with containerization and sandboxed execution

- Experience building or operating high-throughput backend systems

- Hands-on experience building with LLMs, including agent loops, tool calls, and eval harnesses

- Demonstrated ability to own ambiguous problems end-to-end and drive them to a shipped system

- Excellent written and verbal communication skills

## Preferred Qualifications

- Direct experience building RL environments, agentic benchmarks, or eval harnesses

- Familiarity with post-training methods, including RLHF and RLAIF

- Experience designing verifiable reward signals and defending against reward hacking

- Experience with RL training or serving stacks

- Experience with high-scale sandbox or code-execution infrastructure

- Strong observability instincts and experience building internal tools

## Benefits

- Comprehensive health, dental, and vision coverage

- Retirement benefits

- Learning and development stipend

- Generous PTO

- Commuter stipend (may be eligible)

## Salary Range

The base salary range for this full-time position in San Francisco, New York, Seattle is: $252,000-$315,000 USD

## Skills

### Required
- distributed systems
- system design
- data structures
- algorithms
- Python
- containerization
- sandboxed execution
- high-throughput backend systems
- LLMs
- agent loops
- tool calls
- eval harnesses

### Nice to have
- RL environments
- agentic benchmarks
- post-training methods
- RLHF
- RLAIF
- verifiable reward signals
- reward hacking
- RL training or serving stacks
- high-scale sandbox or code-execution infrastructure
- observability

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/scaleai/jobs/4729820005?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
