Consistency is Key: On Data-Efficient Modality Transfer in Speech Translation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | End-to-end approaches to speech translation suffer from data scarcity compared to machine translation (MT). |
| Approach: | They propose a method which combines knowledge distillation and consistency learning to break the dilemma of learning-forgetting. |
| Outcome: | The proposed method outperforms the previous methods on a MuST-C dataset even without additional data. |
Similar Papers
Understanding and Bridging the Modality Gap for Speech Translation (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods to improve end-to-end speech translation (ST) use multitask learning, but there is always a modality gap between ST and MT due to the differences between speech and text. |
| Approach: | They propose a method to bridge the modality gap between ST and MT by leveraging (text) machine translation data. |
| Outcome: | The proposed method bridges the modality gap and achieves significant improvements over baseline in all eight directions. |
Modality Adaption or Regularization? A Case Study on End-to-End Speech Translation (2023.acl-short)
Copied to clipboard
| Challenge: | End-to-end speech translation models have limited training data and are often inefficient due to the inconsistency of length and representation between speech and text. |
| Approach: | They find that the "modality gap" between speech and text data is not a major problem in E2E ST . they decouple the encoder to speech encoder and text encoder, and they find that there is a 'capacity gap' |
| Outcome: | The proposed model achieves 29.0 for en-de and 40.3 for fr on the MuST-C dataset. |
Mutual-Learning Improves End-to-End Speech Translation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to end-to-end speech translation (E2E) models only allow one way knowledge transfer, which is limited by the performance of the teacher model. |
| Approach: | They propose a one-way knowledge transfer paradigm where the MT and ST models are collaboratively trained and considered as peers rather than teacher/student. |
| Outcome: | The proposed model improves the performance of end-to-end speech translation (ST) task by combining knowledge from two models with peer models. |
CKDST: Comprehensively and Effectively Distill Knowledge from Machine Translation to End-to-End Speech Translation (2023.findings-acl)
Copied to clipboard
| Challenge: | End-to-end speech-totext translation (ST) data are limited due to the limited resources. |
| Approach: | They propose a knowledge distillation framework for speech translation that integrates knowledge from machine translation and decouples knowledge from non-target class knowledge. |
| Outcome: | The proposed framework outperforms state-of-the-art models on a benchmark dataset. |
Consistent Transcription and Translation of Speech (2020.tacl-1)
Copied to clipboard
| Challenge: | Existing models that translate without transcribing focus on translation quality, while transcription receives less emphasis. |
| Approach: | They propose a method to evaluate consistency and compare different approaches . they propose 'coupled inference' models that feature a coupled inference procedure can achieve strong consistency. |
| Outcome: | The proposed model is poorly suited to the joint transcription/translation task, but is strong enough to train for consistency. |
Rethinking and Improving Multi-task Learning for End-to-end Speech Translation (2023.emnlp-main)
Copied to clipboard
| Challenge: | auxiliary tasks are highly consistent with end-to-end speech translation (ST) but their effectiveness has not been thoroughly studied. |
| Approach: | They propose an improved multi-task learning approach for the ST task that bridges the modal gap by mitigating the difference in length and representation. |
| Outcome: | The proposed approach achieves state-of-the-art on the MuST-C dataset with 20.8% of training time required by the current SOTA method. |
An Empirical Study of Consistency Regularization for End-to-End Speech-to-Text Translation (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for speech-to-text translation (ST) have achieved impressive supervised and zero-shot performance. |
| Approach: | They propose to use consistency regularization methods to boost end-to-end (E2E) speech-totext translation (ST) by regularizing the intra-modal consistency instead of the modality gap. |
| Outcome: | The proposed training strategies achieve state-of-the-art (SOTA) performance in most translation directions. |
Continual Knowledge Distillation for Neural Machine Translation (2023.acl-long)
Copied to clipboard
| Challenge: | Current parallel corpora are not publicly accessible but trained models are more readily available. |
| Approach: | They propose a method to take advantage of existing translation models to improve one model of interest. |
| Outcome: | The proposed method improves on Chinese-English and German-English datasets and is robust to malicious models. |
MT-PATCHER: Selective and Extendable Knowledge Distillation from Large Language Models for Machine Translation (2024.naacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown their strong ability in the field of machine translation, yet they suffer from high computational cost and latency. |
| Approach: | They propose a framework which transfers knowledge from LLMs to existing MT models in a selective, comprehensive and proactive manner. |
| Outcome: | The proposed framework transfers knowledge from LLMs to existing MT models in a selective, comprehensive and proactive manner. |
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data? (2024.acl-long)
Copied to clipboard
| Challenge: | Existing two-pass direct speech-to-speech translation models require parallel speech data to train, which is challenging to collect. |
| Approach: | They propose a two-pass direct speech-to-speech translation (S2ST) model that decomposes the task into speech- to-text translation (s2TT) and text-tospech (TTS) they propose 'composer' S2ST model that integrates pretrained S2TT and TTS models into a direct S2 ST model. |
| Outcome: | The proposed model integrates pretrained S2TT and TTS models into a direct S2ST model without parallel speech data. |