# Senior Machine Learning Engineer, Voice Agents - EMEA Remote

**Company**: Hugging Face
**Location**: Paris, Île-de-France
**Work arrangement**: remote
**Experience**: senior
**Job type**: Full-time
**Category**: Engineering
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q108943604

**Apply**: https://apply.workable.com/j/9E2A4C02C7?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_e574b5d0-e33

## Description

At Hugging Face, we're on a journey to democratize good AI.

We are building the open voice-agent stack for Hugging Face, and we are looking for a senior engineer to own a large part of it. Two things sit at the centre of this role. The first is speech-to-speech, our open-source library for realtime voice agents. The second is hf-voice, a new product that will let any developer build and deploy voice agents with their Hugging Face account.

Your missions:

- Own the Open-Source Library:

- Take architectural ownership of large parts of speech-to-speech: pipeline design, latency budget, and the reliability of the realtime loop.

- Integrate new ASR, TTS and end-to-end speech models as they land, and keep the abstractions clean while the model landscape keeps moving.

- Review community PRs, triage issues, cut releases, and grow the group of contributors around the project.

- Ship hf-voice:

- Design the developer API and the streaming protocol: session lifecycle, transport (WebSockets/WebRTC), authentication, error semantics, versioning.

- Build the serving side: realtime inference on GPU, concurrency, autoscaling, observability, and cost per session.

- Work with the Hub and inference teams so that a working voice agent is easy to integrate into products and demos.

- Take the product from demo to production: load testing, SLOs, graceful degradation when a model or a network path misbehaves.

- Work in the Open:

- Write the docs, examples and templates that get a developer from zero to a running agent in minutes.

- Support the deployments already relying on the stack, starting with the Reachy Mini fleet.

- Talk about the work publicly if you enjoy it: blog posts, demos, conference talks.

Requirements:

- Senior engineer, able to own a substantial part of an architecture and drive it forward autonomously.

- Experience building developer-facing infrastructure at an AI or developer-tools company: inference APIs, agent infrastructure, or something comparable.

- Substantial open-source contributions to a Python library. Comfortable with async Python and distributed systems, including their failure modes.

- You have shipped something realtime: streaming, WebSockets or WebRTC, audio or video pipelines, live inference.

- Practical experience with LLMs or multimodal models in production. Clear written communication and a habit of collaborating async and in public.

- Motivated by voice and conversational AI.

Bonus points if you have:

- Contributions to a voice-agent framework such as speech-to-speech, pipecat, LiveKit Agents, Vocode or TEN.

- Contributions to llama.cpp or another low-level inference runtime.

- Hands-on work with ASR, TTS or end-to-end speech models, including evaluation of latency and quality trade-offs.

- GPU serving, quantization, or on-device inference experience.

- Audio pipeline knowledge: VAD, echo cancellation, jitter buffers, barge-in and turn detection.

- Experience shipping to embedded or robotics targets.

- A public track record: talks, blog posts, demos.

## Skills

### Required
- Python
- async Python
- distributed systems
- speech-to-speech
- hf-voice
- ASR
- TTS
- end-to-end speech models
- LLMs
- multimodal models
- WebSockets
- WebRTC
- GPU serving
- audio pipelines

### Nice to have
- voice-agent frameworks
- low-level inference runtime
- quantization
- on-device inference
- audio pipeline knowledge
- public track record

---

Source: [Apply at apply.workable.com](https://apply.workable.com/j/9E2A4C02C7?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
