Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS
Researchers have developed Chatterbox-Flash, a text-to-speech model that can generate speech without prior training data. The model uses a block-diffusion decoder to enable parallel token generation within each block while retaining streaming capabilities. To improve performance, the researchers introduced two inference-time techniques: prior-calibrated scoring and an early-decoding schedule. These methods mitigate issues with long-tail token distributions and allow for high-
Researchers have developed Chatterbox-Flash, a text-to-speech model that can generate speech without prior training data. The model uses a block-diffusion decoder to enable parallel token generation within each block while retaining streaming capabilities. To improve performance, the researchers introduced two inference-time techniques: prior-calibrated scoring and an early-decoding schedule. These methods mitigate issues with long-tail token distributions and allow for high-fidelity synthesis comparable to strong baselines. Chatterbox-Flash supports streaming inference with competitive time-to-first-packet and real-time factor.
---
Why it matters: This matters because it enables efficient and high-quality text-to-speech synthesis, which is crucial in applications where real-time speech generation is required, such as voice assistants or live broadcasts.
Source: https://arxiv.org/abs/2605.30748
This article was originally published at: https://arxiv.org/abs/2605.30748