Speech Processing
-
- CONFERENCE (INTERNATIONAL)
- Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
- Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana
- The 25th Annual Conference of the International Speech Communication Association (INTERSPEECH 2024)
- September 04, 2024
-
- CONFERENCE (DOMESTIC)
- トピックモデルを用いた教師なし学習によるHuBERTの意味表現向上
- 前角 高史, Jiatong Shi (カーネギーメロン大学), Xuankai Chang (カーネギーメロン大学), 藤田 悠哉, 渡部 晋治 (カーネギーメロン大学)
- 日本音響学会 2024年秋季研究発表会 (ASJ 2024 autumn)
- September 04, 2024
-
- CONFERENCE (DOMESTIC)
- 離散トークン音声認識におけるドメイン適応の検討
- 石井 敬章, 小松 達也, 藤田 雄介, 藤田 悠哉
- 日本音響学会 2024年秋季研究発表会 (ASJ 2024 autumn)
- September 04, 2024
-
- CONFERENCE (INTERNATIONAL)
- Audio Fingerprinting with Holographic Reduced Representations
- Yusuke Fujita, Tatsuya Komatsu
- The 25th Annual Conference of the International Speech Communication Association (INTERSPEECH 2024)
- September 01, 2024
-
- OTHERS (INTERNATIONAL)
- Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework
- Hokuto Munakata, Ryo Terashima, Yusuke Fujita
- arXiv.org (arXiv)
- June 24, 2024
-
- CONFERENCE (DOMESTIC)
- 「音学シンポジウム2024」開催にあたって
- 大石 康智 (日本電信電話株式会社), 中村 栄太 (九州大学), 大町 基, 森川 大輔 (富山県立大学), 伊藤 信貴 (東京大学), 森 大毅 (宇都宮大学)
- 音学シンポジウム 2024 (第140回MUS・第152回SLP合同研究発表会)
- June 13, 2024
-
- CONFERENCE (INTERNATIONAL)
- Audio Difference Learning for Audio Captioning
- Tatsuya Komatsu, Yusuke Fujita, Kazuya Takeda (Nagoya University), Tomoki Toda (Nagoya University)
- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
- April 14, 2024
-
- CONFERENCE (INTERNATIONAL)
- Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter Sharing
- Brian Yan (Carnegie Mellon University), Xuankai Chang (Carnegie Mellon University), Antonios Anastasopoulos (George Mason University), Yuya Fujita, Shinji Watanabe (Carnegie Mellon University)
- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
- April 14, 2024
-
- CONFERENCE (INTERNATIONAL)
- Enhancing Multilingual TTS with Voice Conversion Based Data Augmentation and Posterior Embedding
- Hyun-Wook Yoon (NAVER Cloud), Jin-Seob Kim (NAVER Cloud), Ryuichi Yamamoto, Ryo Terashima, Chan-Ho Song (NAVER Cloud), Jae-Min Kim (NAVER Cloud), Eunwoo Song (NAVER Cloud)
- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
- April 14, 2024
-
- CONFERENCE (INTERNATIONAL)
- Keep Decoding Parallel With Effective Knowledge Distillation From Language Models To End-To-End Speech Recognisers
- Michael Hentschel (LINE WORKS Corporation), Yuta Nishikawa (Nara Institute of Science and Technology), Tatsuya Komatsu, Yusuke Fujita
- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
- April 14, 2024
-
- CONFERENCE (INTERNATIONAL)
- PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language Descriptions
- Reo Shimizu (Tohoku University), Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, Kentaro Tachibana
- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
- April 14, 2024
-
- OTHERS (INTERNATIONAL)
- LV-CTC: Non-autoregressive ASR with CTC and latent variable models
- Yuya Fujita, Shinji Watanabe (Carnegie Mellon Univ.), Xuankai Chang (Carnegie Mellon Univ.), Takashi Maekaku
- arXiv.org (arXiv)
- March 28, 2024
-
- CONFERENCE (INTERNATIONAL)
- Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study
- Xuankai Chang (Carnegie Mellon University), Brian Yan (Carnegie Mellon University), Kwanghee Choi (Carnegie Mellon University), Jee-Weon Jung (Carnegie Mellon University), Yichen Lu (Carnegie Mellon University), Soumi Maiti (Carnegie Mellon University), Roshan Sharma (Carnegie Mellon University), Jiatong Shi (Carnegie Mellon University), Jinchuan Tian (Carnegie Mellon University), Shinji Watanabe (Carnegie Mellon University), Yuya Fujita, Takashi Maekaku, Pengcheng Guo (Northwestern Polytechnical University), Yao-Fei Cheng (University of Washington), Pavel Denisov (University of Stuttgart), Kohei Saijo (Waseda University), Hsiu-Hsuan Wang (National Taiwan University)
- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
- March 20, 2024
-
- CONFERENCE (INTERNATIONAL)
- Hubertopic: Enhancing Semantic Representation of Hubert Through Self-Supervision Utilizing Topic Model
- Takashi Maekaku, Jiatong Shi (Carnegie Mellon University), Xuankai Chang (Carnegie Mellon University), Yuya Fujita, Shinji Watanabe (Carnegie Mellon University)
- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
- March 20, 2024
-
- CONFERENCE (DOMESTIC)
- 日本語テキストと音楽の対照学習の実験的評価
- 蓮実 拓也, 小松 達也, 藤田 雄介, 二又 航介, 橘 健太郎
- 日本音響学会 2024年春季研究発表会 (ASJ 2024 spring)
- March 07, 2024
-
- CONFERENCE (DOMESTIC)
- 拡散過程と敵対的学習の併用による普遍音声強調
- シャイブラー ロビン, 藤田 雄介, 橘 健太郎
- 日本音響学会 2024年春季研究発表会 (ASJ 2024 spring)
- March 06, 2024
-
- CONFERENCE (DOMESTIC)
- 音声品質と音響環境の潜在変数で条件付けたDenoising Trainingによるノイズロバスト音声変換
- 五十嵐 琢斗 (東京大学), 齋藤 佑樹 (東京大学), 関 健太郎 (東京大学), 高道 慎之介 (東京大学), 山本 龍一, 橘 健太郎, 猿渡 洋 (東京大学)
- 電子情報通信学会/日本音響学会 音声研究会 (IEICE/ASJ-SP)
- February 22, 2024
-
- WORKSHOP (INTERNATIONAL)
- A Comparative Study of Voice Conversion Models with Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 2023
- Ryuichi Yamamoto (Nagoya University / LINE Corp.), Reo Yoneyama (Nagoya University), Lester Phillip Violeta (Nagoya University), Wen-Chin Huang (Nagoya University), Tomoki Toda (Nagoya University)
- The 2023 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU 2023)
- December 19, 2023
-
- CONFERENCE (INTERNATIONAL)
- Domain Adaptation by Data Distribution Matching via Submodularity for Speech Recognition
- Yusuke Shinohara, Shinji Watanabe (Carnegie Mellon University)
- The 2023 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU 2023)
- December 16, 2023
-
- CONFERENCE (INTERNATIONAL)
- LV-CTC: Non-autoregressive ASR with CTC and Latent Variable Models
- Yuya Fujita, Shinji Watanabe (Carnegie Mellon Univ.), Xuankai Chang (Carnegie Mellon Univ.), Takashi Maekaku
- The 2023 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU 2023)
- December 16, 2023