Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing speech translation approaches often overlook the transfer of speech patterns, leading to mismatches with source speech and limiting their suitability for dubbing applications. |
| Approach: | They propose a diffusion-based speech-to-unit translation model with explicit duration control that enables time-aligned translation. |
| Outcome: | The proposed system preserves key characteristics such as duration, speaker identity, and speaking speed while maintaining key characteristics. |
Similar Papers
Textless Speech-to-Speech Translation on Real Data (2022.naacl-main)
Copied to clipboard
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, Wei-Ning Hsu
| Challenge: | Existing text-based speech-to-speech translation systems rely on cascaded approach . text-to text translation systems require text generation and a single input to generate output . |
| Approach: | They propose a textless speech-to-speech translation system that can translate speech from one language into another without the need of text data. |
| Outcome: | The proposed system can translate speech from one language into another without text data. |
Direct Speech-to-Speech Translation With Discrete Units (2022.acl-long)
Copied to clipboard
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, Wei-Ning Hsu
| Challenge: | Existing direct speech-to-speech translation models rely on text generation as an intermediate step. |
| Approach: | They propose a direct speech-to-speech translation model that translates speech from one language to another without relying on intermediate text generation. |
| Outcome: | The proposed model produces 6.7 BLEUs in the Fisher Spanish-English dataset when trained without any text transcripts and with text supervision. |
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation (2025.findings-acl)
Copied to clipboard
Yuhao Zhang, Xiangnan Ma, Kaiqi Kou, Peizhuo Liu, Weiqiao Shan, Benyou Wang, Tong Xiao, Yuxin Huang, Zhengtao Yu, JingBo Zhu
| Challenge: | Existing textless speech-to-speech translation models have two main challenges: 1) learning cross-modal features and 2) learning alignment of difference languages in long sequences. |
| Approach: | They propose a unit language to overcome two main modeling challenges . they propose task prompt modeling to utilize the unit language in guiding the modeling process. |
| Outcome: | The proposed language improves over a strong baseline and achieves comparable performance to models trained with text. |
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods to translate spoken utterances from one language to another are unable to preserve speaker timbre of source speech. |
| Approach: | They propose a pipeline with style-transfer capability on the basis of self-supervised speech representations and codec units. |
| Outcome: | The proposed model achieves zero-shot cross-lingual style transfer on previously unseen source languages. |
Speech-to-Speech Translation for a Real-world Unwritten Language (2023.findings-acl)
Copied to clipboard
Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du, Justine Kao, Yu-An Chung, Paden Tomasello, Paul-Ambroise Duquenne, Holger Schwenk, Hongyu Gong, Hirofumi Inaguma, Sravya Popuri, Changhan Wang, Juan Pino, Wei-Ning Hsu, Ann Lee
| Challenge: | a new study examines speech-to-speech translation (S2ST) that translates speech from one language into another . the research area for unwritten languages remains a research area with little exploration due to the lack of training data. |
| Approach: | They propose a system that translates speech from one language into another . they use Taiwanese Hokkien as an example of an unwritten language . |
| Outcome: | The proposed system can be used to train models in languages without standard writing systems. |
DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing models for speech generation are not efficient due to low information density of speech data. |
| Approach: | They propose a method to integrate discrete diffusion models into speech generation tasks . they propose to apply diffusion forward process while employing diffusion backward process . |
| Outcome: | The proposed model achieves comparable results to the auto-regressive baselines with significantly fewer decoding steps (50 steps). |
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data? (2024.acl-long)
Copied to clipboard
| Challenge: | Existing two-pass direct speech-to-speech translation models require parallel speech data to train, which is challenging to collect. |
| Approach: | They propose a two-pass direct speech-to-speech translation (S2ST) model that decomposes the task into speech- to-text translation (s2TT) and text-tospech (TTS) they propose 'composer' S2ST model that integrates pretrained S2TT and TTS models into a direct S2 ST model. |
| Outcome: | The proposed model integrates pretrained S2TT and TTS models into a direct S2ST model without parallel speech data. |
Textless Speech-to-Speech Translation With Limited Parallel Data (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing speech-to-speech translation models either leverage text as an intermediate step or require hundreds of hours of parallel speech data. |
| Approach: | They propose a framework for training textless S2ST models that require dozens of hours of parallel speech data. |
| Outcome: | The proposed model achieves reasonable performance on three domains with single-speaker synthesized speech. |
Simple and Effective Unsupervised Speech Translation (2023.acl-long)
Copied to clipboard
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, Juan Pino
| Challenge: | Existing methods to train speech models without labeled data are limited for most languages. |
| Approach: | They propose a pipeline approach to build speech translation systems without labeled data by leveraging recent advances in unsupervised speech recognition, machine translation and speech synthesis. |
| Outcome: | The proposed approach outperforms the state-of-the-art in unsupervised speech recognition by 3.2 BLEU on the Libri-Trans benchmark and the best supervised end-to-end models from only two years ago by an average of 5.0 BLUE over five X-En directions. |
Fluent and Low-latency Simultaneous Speech-to-Speech Translation with Self-adaptive Training (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Current approaches to simultaneous speech-to-speech translation accumulate more and more latencies in later sentences when the speaker talks faster. |
| Approach: | They propose a method which generates more fluent target speech latency than the baseline . they propose to use self-adaptive translation to adjust the length of translations to accommodate different source speech rates. |
| Outcome: | Xiong et al., 2019) show that the proposed method generates more fluent target speech latency than baseline . authors say it provides more natural communication process than speech-to-text translation . xiong and colleagues say the proposed technique is more efficient than current approaches . |