# Member of Technical Staff - RL Inference

**Company**: xAI
**Location**: Palo Alto, CA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q120599684

**Apply**: https://job-boards.greenhouse.io/xai/jobs/5180223007?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_a9208813-5ef

## Description

## About the Role

We're seeking an engineer to join our RL infrastructure team, focusing on low-precision RL training and inference.

## Responsibilities

- Design and optimize our inference stack for various RL workloads, from small-scale ablations to production training runs.

- Analyze, profile, and address performance bottlenecks in large-scale RL systems.

- Collaborate with the modelling team to implement novel RL techniques and algorithms efficiently.

## Basic Qualifications

- Experience building, debugging, and optimizing large-scale distributed systems.

- Experience with LLM inference.

- Proficiency in languages like Python, C++, and/or Rust; frameworks such as PyTorch, Jax, CUDA.

- Willingness to tackle complex problems across all levels of the stack.

## Preferred Skills and Experience

- Strong knowledge of quantization and numerics in LLM inference and training.

- Experience developing inference engines, e.g., SGLang, vLLM.

## Skills

### Required
- large-scale distributed systems
- LLM inference
- Python
- C++
- Rust
- PyTorch
- Jax
- CUDA

### Nice to have
- quantization
- numerics in LLM inference and training
- inference engines
- SGLang
- vLLM

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/xai/jobs/5180223007?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
