Description
About the Role
We're seeking an engineer to join our RL infrastructure team, focusing on low-precision RL training and inference.
Responsibilities
- Design and optimize our inference stack for various RL workloads, from small-scale ablations to production training runs.
- Analyze, profile, and address performance bottlenecks in large-scale RL systems.
- Collaborate with the modelling team to implement novel RL techniques and algorithms efficiently.
Basic Qualifications
- Experience building, debugging, and optimizing large-scale distributed systems.
- Experience with LLM inference.
- Proficiency in languages like Python, C++, and/or Rust; frameworks such as PyTorch, Jax, CUDA.
- Willingness to tackle complex problems across all levels of the stack.
Preferred Skills and Experience
- Strong knowledge of quantization and numerics in LLM inference and training.
- Experience developing inference engines, e.g., SGLang, vLLM.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/xai/jobs/5180223007