Publications

カンファレンス (国際) Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

Jinchuan Tian (Carnegie Mellon University), Haoran Wang (Carnegie Mellon University), Siddhant Arora (Carnegie Mellon University), Takashi Maekaku, Keita Goto, Jin Sakuma, Yusuke Shinohara, Chao-Han Huck Yang (NVIDIA Research), Shinji Watanabe (Carnegie Mellon University)

The 27th Annual Conference of the International Speech Communication Association (INTERSPEECH 2026)

2026.9.27

Classical TTS systems typically rely on rigid input formats and predefined metadata slots, limiting their ability to fulfill flexible user requirements. This paper introduces Bagpiper-TTS, a universal speech synthesis system that deals with diverse natural language user requests. Given a natural language prompt, Bagpiper-TTS first reasons over the user's intent to derive a rich caption, i.e., a comprehensive textual blueprint encompassing both transcription and nuanced metadata. Subsequently, this caption guides the synthesis of the target speech. Our model inherently supports a broad spectrum of tasks besides classical TTS applications, including multi-talker, intent-to-speech, role-play synthesis, singing voice synthesis, and more. Experimental results demonstrate that Bagpiper-TTS achieves an 1.7% Word Error Rate (WER) on the Seed-TTS-Eval benchmark and match the performance of dedicated models in both LLM-as-a-judge and human subjective evaluations across multiple applications.

Paper : Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis新しいタブまたはウィンドウで開く (外部サイト)