Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions
Researchers have developed Poly-InstructTTS, a system for generating expressive speech from open-ended instructions. The approach uses in-the-wild audiovisual data and a multi-modal pipeline to learn fine-grained emotions and styles. A prompt-free GPT model is used with attribute-based thinking tokens, followed by a flow-matching module that injects timbre from a reference audio. Experiments show strong performance in instruction adherence and expressiveness.
Researchers have developed Poly-InstructTTS, a system for generating expressive speech from open-ended instructions. The approach uses in-the-wild audiovisual data and a multi-modal pipeline to learn fine-grained emotions and styles. A prompt-free GPT model is used with attribute-based thinking tokens, followed by a flow-matching module that injects timbre from a reference audio. Experiments show strong performance in instruction adherence and expressiveness.
---
Why it matters: This matters because it addresses the challenge of controlling fine-grained expression in text-to-speech models using natural-language instructions, which is essential for applications like voice assistants and virtual characters.
Source: https://arxiv.org/abs/2608.20387
This article was originally published at: https://arxiv.org/abs/2608.20387