MuST-C: a Multilingual Speech Translation Corpus (N19-1)

Copied to clipboard

Challenge: Current research on spoken language translation (SLT) has to confront the scarcity of sizeable and publicly available training corpora.
Approach: They propose a multilingual speech translation corpus that will facilitate the training of end-to-end systems for SLT from English into 8 languages.
Outcome: The proposed multilingual speech translation corpus will facilitate the training of end-to-end systems for spoken language translation from English into 8 languages.

Similar Papers

Simul-MuST-C: Simultaneous Multilingual Speech Translation Corpus Using Large Language Model (2024.emnlp-main)

Copied to clipboard

Challenge: Simultaneous speech translation (SiST) begins translating before the entire source input is received.
Approach: They propose a dataset that rearranges sentences into segmented monotonic data for simultaneous speech translation using the Large Language Model.
Outcome: The proposed dataset improves quality and latency in siST translations by rearranging sentences into segmented monotonic data.
MuST-Cinema: a Speech-to-Subtitles corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for subtitling are laborious and costly, says aaron sanchez . he says the current methods are laboriously complex and require manual work .
Approach: They propose to use TED subtitles to build a multilingual speech translation corpus . they propose to annotate existing subtitling corpora with subtitle breaks .
Outcome: The proposed model can be used to segment sentences into subtitles and reduces human work . the proposed model reduces the time and cost of human subtitling tasks .
Rethinking and Improving Multi-task Learning for End-to-end Speech Translation (2023.emnlp-main)

Copied to clipboard

Challenge: auxiliary tasks are highly consistent with end-to-end speech translation (ST) but their effectiveness has not been thoroughly studied.
Approach: They propose an improved multi-task learning approach for the ST task that bridges the modal gap by mitigating the difference in length and representation.
Outcome: The proposed approach achieves state-of-the-art on the MuST-C dataset with 20.8% of training time required by the current SOTA method.
CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets involve language pairs with English as source language, are low resource or lack labeled data.
Approach: They propose a multilingual speech-to-text translation corpus from 11 languages into English . they provide empirical evidence of the quality of the data and provide initial benchmarks .
Outcome: The proposed model is the first end-to-end multilingual model for spoken language translation.
MUST&P-SRL: Multi-lingual and Unified Syllabification in Text and Phonetic Domains for Speech Representation Learning (2023.emnlp-industry)

Copied to clipboard

Challenge: Modern speech technologies have moved towards end-to-end models that constitutes black box systems that do not allow for explainability of the prediction or decisions.
Approach: They propose a method for linguistic feature extraction that uses phonetic transcriptions and a forced alignment tool to extract phonetic features.
Outcome: The proposed method is compatible with a forced-alignment tool, the Montreal Forced Aligner.
CVSS Corpus and Massively Multilingual Speech-to-Speech Translation (2022.lrec-1)

Copied to clipboard

Challenge: Existing work on speech-to-speech translation (S2ST) systems rely on text representation, but they are text-centric.
Approach: They introduce a massively multilingual-to-English speech-tospeech translation corpus . they synthesize the translation text from the Common Voice speech corpus and CoVoST 2 into English .
Outcome: The proposed corpus outperforms existing models on CoVoST 2 by 5.8 BLEU . the proposed model outperformed the previous state-of-the-art model without extra data .
Khan Academy Corpus: A Multilingual Corpus of Khan Academy Lectures (2024.lrec-main)

Copied to clipboard

Challenge: a dataset of 10122 hours in 87394 recordings is presented in a new journal . 43% of recordings have human-written subtitles, covering a total of 137 languages.
Approach: They present a Khan Academy corpus with 10122 hours in 87394 recordings . 43% of recordings have human-written subtitles, and 137 languages are included .
Outcome: The dataset can be used to train multilingual speech recognition and translation models.
Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation (L18-1)

Copied to clipboard

Challenge: Recent work in spoken language translation (SLT) has attempted to build end-to-end speech-totext translation without using source language transcription during learning or decoding.
Approach: They propose to augment an existing (monolingual) corpus: LibriSpeech.
Outcome: The proposed corpus is derived from read audiobooks from the LibriVox project and has been carefully segmented and aligned.
The Multilingual Microblog Translation Corpus: Improving and Evaluating Translation of User-Generated Text (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of over 200,000 microblog translations supports translation of thirteen languages into English . large collections of parallel text, or bitext, are increasingly available in many languages .
Approach: They propose a corpus of over 200,000 microblog posts that supports translation of thirteen languages into English.
Outcome: The proposed corpus contains over 200,000 translations of microblog posts in 13 languages . fine-tuning showed significant improvements in translation quality .
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations