For voice AI, latency is a essential parameter. Builders have made super progress in mannequin high quality, however the person expertise continues to be typically restricted by response occasions. Hugging Face and Cerebras are altering that have. Immediately, we show what turns into potential when an open, modular voice AI structure is paired with industry-leading inference velocity.
The result’s a speech-to-speech expertise that feels dramatically extra pure. As a substitute of ready for an AI to reply, conversations stream with the responsiveness customers count on from human interplay.
Structure: an Open, Cascaded Speech-to-Speech stack
The demo is constructed as a real-time speech-to-speech pipeline. Every a part of the system is modular, open, and replaceable, making it simple for builders to adapt the stack for various assistants, robots, merchandise, or analysis tasks.
This creates a completely open speech-to-speech loop:
Speech enter
-> speech recognition with Nvidia’s Parakeet
-> Gemma 4 VLM inference on Cerebras
-> text-to-speech with Alibaba’s Qwen3TTS
-> spoken response
The structure brings collectively the power of the open-source AI ecosystem: Cerebras for quick inference, Google DeepMind’s Gemma 4 31B for the language mannequin, and Qwen for text-to-speech. Each layer will be inspected, modified, and prolonged by the builders
Cerebras and Hugging Face Partnership
Immediately, some manufacturing methods see an inexpensive median latency whereas nonetheless experiencing irritating multi-second delays on the P95. These delays turn out to be much more noticeable when device calls or multimodal steps require a number of turns.
Cerebras helps remedy some of the essential bottlenecks within the stack: the language-model response time. By making inference dramatically sooner and extra steady, Cerebras permits the remainder of the Hugging Face pipeline to shine.
That stability is particularly essential on the lengthy tail. Many methods can ship acceptable median response occasions, however occasional gradual responses nonetheless make conversations really feel unreliable.
Constructed for real-world interplay
This similar Hugging Face speech-to-speech pipeline already powers Reachy Mini robots, with greater than 9,000 robots within the wild. For robots, voice assistants, and embodied AI, responsiveness isn’t a beauty enchancment. It’s what makes the interplay really feel alive.
The motivation to make use of Cerebras is subsequently not merely price discount. It’s low latency, predictable efficiency, and the flexibility to create real-time experiences that really feel pure at scale.
This collaboration displays a shared perception that the way forward for AI shall be each open and performant. Open-source fashions, open infrastructure, and breakthrough inference velocity collectively create a basis for the subsequent era of conversational AI.
We invite builders to discover the demo, experiment with the code, and assist form what comes subsequent for real-time voice AI.
Demo: Hugging Face House
Repository: huggingface/speech-to-speech
![[2602.11354] ReplicatorBench: Benchmarking LLM Brokers for Replicability in Social and Behavioral Sciences [2602.11354] ReplicatorBench: Benchmarking LLM Brokers for Replicability in Social and Behavioral Sciences](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)