Challenge: XEUS is a cross-lingual encoder for universal speech that can be trained on 1 million hours of data across 4057 languages.
Approach: They propose a Cross-lingual Encoder for Universal Speech that can be trained on 1 million hours of data across 4057 languages and a newly created corpus of 7400+ hours from 4057 .
Outcome: The proposed model outperforms state-of-the-art models on several benchmarks and outperfies MMS 1B and w2v-BERT 2.0 v2 by 0.8% and 4.4% respectively.

Similar Papers

Self-supervised Representation Learning for Speech Processing (2022.naacl-tutorials)

Copied to clipboard

Challenge: Self-supervised representation learning (SSL) uses proxy supervised learning tasks to obtain training data from unlabeled corpora.
Approach: They propose to survey the latest SSL techniques, tools, datasets, and performance achievement in speech processing to scale up current machine learning technologies.
Outcome: The proposed tutorial is highly relevant to the special theme of ACL about language diversity.
Audiocite.net : A Large Spoken Read Dataset in French (2024.lrec-main)

Copied to clipboard

Challenge: Existing self-supervised learning methods for speech processing have proved difficult to apply to French due to the scarcity of large speech datasets.
Approach: They present a corpus of 6,682 hours of audiobooks from 130 readers . they describe the creation process and final statistics of the corpus .
Outcome: The proposed model based on the audiocite.net corpus, which contains 6,682 hours of audiobooks, was able to perform in 14k version.
SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation (2026.acl-long)

Copied to clipboard

Challenge: Empirically, SpidR-Adapt achieves rapid gains in phonemic discriminability and downstream spoken language modeling scores . current self-supervised learning models require thousands of hours of training data to learn meaningful linguistic representations.
Approach: They propose a bi-level optimization framework for rapid adaptation of speech units to new languages using minimal unlabeled data.
Outcome: The proposed model achieves rapid gains in phonemic discriminability and spoken language modeling scores . it surpasses in-domain toplines after training on less than 1h of target-language audio .
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception (2024.acl-long)

Copied to clipboard

Challenge: Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments.
Approach: They propose a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages.
Outcome: The proposed model outperforms the previous state-of-the-art by 18.5% WER and 4.7 BLEU on downstream audio-visual speech recognition and translation tasks.
Unsupervised Cross-lingual Representation Learning at Scale (2020.acl-main)

Copied to clipboard

Challenge: Pretraining multilingual language models at scale leads to performance gains for cross-lingual transfer tasks.
Approach: They present a transformer-based multilingual masked language model pre-trained on 100 languages . they show that pretraining multilingual models at scale leads to significant performance gains .
Outcome: The proposed model outperforms multilingual BERT (mBERT) on cross-lingual benchmarks.
Speech-to-Speech Translation for a Real-world Unwritten Language (2023.findings-acl)

Copied to clipboard

Challenge: a new study examines speech-to-speech translation (S2ST) that translates speech from one language into another . the research area for unwritten languages remains a research area with little exploration due to the lack of training data.
Approach: They propose a system that translates speech from one language into another . they use Taiwanese Hokkien as an example of an unwritten language .
Outcome: The proposed system can be used to train models in languages without standard writing systems.
Evaluating Self-Supervised Speech Representations for Indigenous American Languages (2024.lrec-main)

Copied to clipboard

Challenge: a recent study focused on the use of self-supervised learning to learn speech representations for indigenous languages . aaron e. scott: the vast linguistic diversity represented by indigenous languages remains unexplored . by expanding the scope of language processing to include indigenous languages, we can foster linguistic inclusivity, he says .
Approach: They benchmark the efficacy of large-scale self-supervised learning models on indigenous American languages.
Outcome: The proposed model can generalize to real-world data, showing strong performance . evaluators found that the model performed better than monolingual models on indigenous languages .
Beyond English-Centric Bitexts for Better Multilingual Language Representation Learning (2023.acl-long)

Copied to clipboard

Challenge: XY-LENT: X-Y bitext enhanced Language ENcodings achieves state-of-the-art performance over 5 cross-lingual tasks within all model size bands.
Approach: They propose a method for building multilingual representation models that are competitive with existing models and more parameter efficient.
Outcome: The proposed model outperforms XLM-R XXL and is 5x and 6x smaller respectively.
LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback (2024.findings-acl)

Copied to clipboard

Challenge: Recent multilingual models support limited number of human languages due to lack of training data for low resource languages.
Approach: They propose a multilingual multilingual LLM that scales to 100 languages . they use a human feedback dataset and a data set to perform multilingual instruction tuning .
Outcome: The proposed model outperforms its peers on five multilingual benchmarks.
LaRA: Large Rank Adaptation for Speech and Text Cross-Modal Learning in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to integrate speech and text capabilities into large language models (LLMs) require significantly larger ranks comparable to the pretrained weights to accommodate the complexities of speech-text cross-modality learning.
Approach: They propose a large-rank adaptive approach for cross-modal integration of speech and text into large language models (LLMs) it uses a Hi-Fi vocoder to synthesize speech waveforms from the generated speech units.
Outcome: The proposed model can be extended to other cross-modal applications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations