Description
We are looking for a Research Engineer to join the research team at ElevenLabs, focused on deploying and optimizing our frontier AI models in production.
The quality of our models only matters if they can be served fast, reliably, and at scale. You will own the systems that turn research breakthroughs into real-time products used by millions.
Responsibilities:
- Deploying state-of-the-art models to production and owning the path from research checkpoint to serving infrastructure.
- Optimizing inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.
- Building and tuning high-performance serving systems for real-time, streaming workloads where every millisecond matters.
- Creating tooling and infrastructure that lets researchers ship new models to production quickly, safely, and with confidence in their performance characteristics.
Requirements:
- Experience deploying and serving ML models in production, ideally for latency-sensitive or real-time applications.
- Strong engineering skills in GPU programming and inference optimization (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang).
- The capacity to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack, from model architecture to kernels to orchestration, and to build the tooling to measure it.
What we offer:
- Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.
- Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.
- Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.
- Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.
- Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy.
- Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://elevenlabs.io/pl/careers/2d7f9a7c-a9e6-4877-bb38-34e4d989054c/research-engineer-inference