Description
At Hugging Face, we're on a journey to democratize good AI.
We are building the open voice-agent stack for Hugging Face, and we are looking for a senior engineer to own a large part of it. Two things sit at the centre of this role. The first is speech-to-speech, our open-source library for realtime voice agents. The second is hf-voice, a new product that will let any developer build and deploy voice agents with their Hugging Face account.
Your missions:
- Own the Open-Source Library:
- Take architectural ownership of large parts of speech-to-speech: pipeline design, latency budget, and the reliability of the realtime loop.
- Integrate new ASR, TTS and end-to-end speech models as they land, and keep the abstractions clean while the model landscape keeps moving.
- Review community PRs, triage issues, cut releases, and grow the group of contributors around the project.
- Ship hf-voice:
- Design the developer API and the streaming protocol: session lifecycle, transport (WebSockets/WebRTC), authentication, error semantics, versioning.
- Build the serving side: realtime inference on GPU, concurrency, autoscaling, observability, and cost per session.
- Work with the Hub and inference teams so that a working voice agent is easy to integrate into products and demos.
- Take the product from demo to production: load testing, SLOs, graceful degradation when a model or a network path misbehaves.
- Work in the Open:
- Write the docs, examples and templates that get a developer from zero to a running agent in minutes.
- Support the deployments already relying on the stack, starting with the Reachy Mini fleet.
- Talk about the work publicly if you enjoy it: blog posts, demos, conference talks.
Requirements:
- Senior engineer, able to own a substantial part of an architecture and drive it forward autonomously.
- Experience building developer-facing infrastructure at an AI or developer-tools company: inference APIs, agent infrastructure, or something comparable.
- Substantial open-source contributions to a Python library. Comfortable with async Python and distributed systems, including their failure modes.
- You have shipped something realtime: streaming, WebSockets or WebRTC, audio or video pipelines, live inference.
- Practical experience with LLMs or multimodal models in production. Clear written communication and a habit of collaborating async and in public.
- Motivated by voice and conversational AI.
Bonus points if you have:
- Contributions to a voice-agent framework such as speech-to-speech, pipecat, LiveKit Agents, Vocode or TEN.
- Contributions to llama.cpp or another low-level inference runtime.
- Hands-on work with ASR, TTS or end-to-end speech models, including evaluation of latency and quality trade-offs.
- GPU serving, quantization, or on-device inference experience.
- Audio pipeline knowledge: VAD, echo cancellation, jitter buffers, barge-in and turn detection.
- Experience shipping to embedded or robotics targets.
- A public track record: talks, blog posts, demos.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://apply.workable.com/j/9E2A4C02C7