Papers with VoxPopuli

5 papers
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation (2021.acl-long)

Copied to clipboard

Challenge: VoxPopuli provides 400K hours of unlabeled speech data in 23 languages . large amounts of multilingual audio data are needed to achieve similar progress for multilingual ASR and ST.
Approach: They propose a large-scale multilingual corpus that provides 400K hours of unlabeled speech data in 23 languages.
Outcome: The proposed corpus provides 400K hours of unlabeled speech data in 23 languages and 1.8K hours transcribed speeches in 15 languages and their aligned oral interpretations into 15 target languages totaling 17.3K hours.
Group Fairness in Multilingual Speech Recognition Models (2024.findings-naacl)

Copied to clipboard

Challenge: a new study evaluates the performance disparities of ASR models across languages and demographics . a large amount of data is required to mitigate performance disparity, but this is computationally expensive .
Approach: They evaluate the performance disparity of ASR models using a multilingual dataset . they find that model size correlates logarithmically with worst-case performance disparities .
Outcome: The proposed models exhibit significant performance disparities across binary genders for adolescents.
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning (2025.acl-long)

Copied to clipboard

Challenge: Despite their growing importance, the quality of these datasets remains under-researched.
Approach: They propose guidelines and recommendations to address quality issues in future dataset development . they find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages .
Outcome: The results highlight the need for proactive language planning and enhanced data quality control in the process of automatic speech recognition dataset creation.
Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents (2025.emnlp-main)

Copied to clipboard

Challenge: Speech Vecalign is a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions.
Approach: They propose a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions.
Outcome: The proposed method outperforms SpeechMatrix models on 3,000 hours of unlabeled speech documents and produces longer speech-to-speech alignments.
SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations (2023.acl-long)

Copied to clipboard

Challenge: SpeechMatrix is a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings.
Approach: They present a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings.
Outcome: The proposed model can train bilingual models on 136 language pairs with 418 thousand hours of speech.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations