Challenge: Recent expressive speech-to-speech translation systems have achieved impressive expressivity preservation performances by cascading unit-to speech (U2S) generator to the speech- to-unit translation model.
Approach: They propose a textless acoustic model with a self-supervised distillation strategy for noise-robust expressive speech-to-speech translation (S2ST) They aim to address this limitation by incorporating a distillation with no label (DINO) self-controlled training strategy into the model’s pretraining process.
Outcome: The proposed model significantly improved the expressive speech-to-speech translation system in noisy environments while maintaining competitive performance in clean environments.

Similar Papers

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation (2023.acl-long)

Copied to clipboard

Challenge: Existing models for speech-to-speech translation suffer from distinct degradation in noisy environments and fail to translate visual speech.
Approach: They propose a text-based audio-visual speech-to-speech translation model that integrates visual information with audio-only data to improve system robustness.
Outcome: The proposed model outperforms models trained on audio-only corpus in two languages . it also improves with low-resource audio-visual data, compared with baselines .
Textless Speech-to-Speech Translation on Real Data (2022.naacl-main)

Copied to clipboard

Challenge: Existing text-based speech-to-speech translation systems rely on cascaded approach . text-to text translation systems require text generation and a single input to generate output .
Approach: They propose a textless speech-to-speech translation system that can translate speech from one language into another without the need of text data.
Outcome: The proposed system can translate speech from one language into another without text data.
CTC-based Non-autoregressive Textless Speech-to-Speech Translation (2024.findings-acl)

Copied to clipboard

Challenge: Existing direct speech-to-speech translation models require text supervision during training, which is not feasible for numerous unwritten languages.
Approach: They propose a non-autoregressive (NAR) model that generates discrete units from the source speech and employs a unit-based vocoder to synthesize the target.
Outcome: The proposed model achieves translation quality comparable to the autoregressive model while preserving up to 26.81 decoding speedup.
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing speech translation approaches often overlook the transfer of speech patterns, leading to mismatches with source speech and limiting their suitability for dubbing applications.
Approach: They propose a diffusion-based speech-to-unit translation model with explicit duration control that enables time-aligned translation.
Outcome: The proposed system preserves key characteristics such as duration, speaker identity, and speaking speed while maintaining key characteristics.
Speech-to-Speech Translation for a Real-world Unwritten Language (2023.findings-acl)

Copied to clipboard

Challenge: a new study examines speech-to-speech translation (S2ST) that translates speech from one language into another . the research area for unwritten languages remains a research area with little exploration due to the lack of training data.
Approach: They propose a system that translates speech from one language into another . they use Taiwanese Hokkien as an example of an unwritten language .
Outcome: The proposed system can be used to train models in languages without standard writing systems.
ParrotTTS: Text-to-speech synthesis exploiting disentangled self-supervised representations (2024.findings-eacl)

Copied to clipboard

Challenge: ParrotTTS can train a multi-speaker variant using transcripts from a single speaker in low resource setup and generalizes to languages not seen while training the self-supervised backbone.
Approach: They propose a modular text-to-speech synthesis model that can train a multi-speaker variant using transcripts from a single speaker.
Outcome: The proposed model outperforms state-of-the-art multi-lingual text-to-speech models using only a fraction of paired data as latter.
Direct Speech-to-Speech Translation With Discrete Units (2022.acl-long)

Copied to clipboard

Challenge: Existing direct speech-to-speech translation models rely on text generation as an intermediate step.
Approach: They propose a direct speech-to-speech translation model that translates speech from one language to another without relying on intermediate text generation.
Outcome: The proposed model produces 6.7 BLEUs in the Fisher Spanish-English dataset when trained without any text transcripts and with text supervision.
Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units (2023.findings-emnlp)

Copied to clipboard

Challenge: DISSC is a lightweight voice conversion method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner.
Approach: They propose a method that converts rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner.
Outcome: The proposed method outperforms baseline methods on quantitative and qualitative evaluations.
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer (2024.acl-srw)

Copied to clipboard

Challenge: Existing methods to translate spoken utterances from one language to another are unable to preserve speaker timbre of source speech.
Approach: They propose a pipeline with style-transfer capability on the basis of self-supervised speech representations and codec units.
Outcome: The proposed model achieves zero-shot cross-lingual style transfer on previously unseen source languages.
DM-Codec: Distilling Multimodal Representations for Speech Tokenization (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing speech tokenization models lack contextual representations for speech synthesis . absence of contextual representation results in elevated WER and WIL scores .
Approach: They propose a language model-guided distillation method that incorporates contextual information into a comprehensive speech tokenizer.
Outcome: The proposed method outperforms state-of-the-art tokenization models in reducing WER and WIL scores.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations