Papers with VoxPopuli
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation (2021.acl-long)
Copied to clipboard
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, Emmanuel Dupoux
| Challenge: | VoxPopuli provides 400K hours of unlabeled speech data in 23 languages . large amounts of multilingual audio data are needed to achieve similar progress for multilingual ASR and ST. |
| Approach: | They propose a large-scale multilingual corpus that provides 400K hours of unlabeled speech data in 23 languages. |
| Outcome: | The proposed corpus provides 400K hours of unlabeled speech data in 23 languages and 1.8K hours transcribed speeches in 15 languages and their aligned oral interpretations into 15 target languages totaling 17.3K hours. |
Group Fairness in Multilingual Speech Recognition Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | a new study evaluates the performance disparities of ASR models across languages and demographics . a large amount of data is required to mitigate performance disparity, but this is computationally expensive . |
| Approach: | They evaluate the performance disparity of ASR models using a multilingual dataset . they find that model size correlates logarithmically with worst-case performance disparities . |
| Outcome: | The proposed models exhibit significant performance disparities across binary genders for adolescents. |
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning (2025.acl-long)
Copied to clipboard
| Challenge: | Despite their growing importance, the quality of these datasets remains under-researched. |
| Approach: | They propose guidelines and recommendations to address quality issues in future dataset development . they find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages . |
| Outcome: | The results highlight the need for proactive language planning and enhanced data quality control in the process of automatic speech recognition dataset creation. |
Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents (2025.emnlp-main)
Copied to clipboard
| Challenge: | Speech Vecalign is a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions. |
| Approach: | They propose a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions. |
| Outcome: | The proposed method outperforms SpeechMatrix models on 3,000 hours of unlabeled speech documents and produces longer speech-to-speech alignments. |
SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations (2023.acl-long)
Copied to clipboard
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk
| Challenge: | SpeechMatrix is a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. |
| Approach: | They present a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. |
| Outcome: | The proposed model can train bilingual models on 136 language pairs with 418 thousand hours of speech. |