Challenge: Recent advances in automatic speech recognition (ASR) systems have been criticized for high acoustic variability and limited amount of available training data.
Approach: They propose a two-step training strategy that uses multilingual learning followed by language-specific transfer learning to generalize children's speech.
Outcome: The proposed training strategy outperforms single language training and multilingual and transfer learning alone in English.

Similar Papers

A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models.
Approach: They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages.
Outcome: The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data.
Is Child-Directed Speech Effective Training Data for Language Models? (2024.emnlp-main)

Copied to clipboard

Challenge: High-performing language models are typically trained on hundreds of billions of words, but human learners use language fluently after far less training data.
Approach: They train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset.
Outcome: The proposed models show that child language input is not valuable for training language models.
Evaluating and Improving Child-Directed Automatic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that adult speech recognition systems are lagging behind child models due to the fact that children's vocal tracts are smaller than adults .
Approach: They evaluate a model that trains on adult data and apply additional tuning to varied amounts of child speech data to improve child-directed speech recognition.
Outcome: The proposed model improves over baseline models using child data and small amounts of child audio data.
Learning from Child-directed Speech in Two-language Scenarios: A French-English Case-Study (2026.findings-eacl)

Copied to clipboard

Challenge: a systematic study of compact language models with limited computational resources is challenging for many research contexts and real-world applications.
Approach: They extend BabyBERTa to English-French scenarios under strictly sizematched data conditions.
Outcome: The proposed model extends to English-French scenarios under sizematched data conditions . the results show context-dependent effects of multilingual training .
Error-preserving Automatic Speech Recognition of Young English Learners’ Language (2024.acl-long)

Copied to clipboard

Challenge: State-of-the-art speech recognition models are often trained on adult read-aloud data by native speakers and do not transfer well to young language learners’ speech.
Approach: They propose to use an automated speech recognition module to train language learners' speaking skills on spontaneous speech by young language learners.
Outcome: The proposed model improves on 85 hours of English audio spoken by Swiss learners and preserves their mistakes.
The Less the Merrier? Investigating Language Representation in Multilingual Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Multilingual models can be used to integrate multiple languages into one model and use cross-language transfer learning to improve performance for different NLP tasks.
Approach: They propose to include languages in popular multilingual models and to use cross-language transfer learning to improve performance for different NLP tasks.
Outcome: The proposed models perform better on downstream tasks for seen and unseen languages than community-centered models for low-resource languages.
Learning to Understand Child-directed and Adult-directed Speech (2020.acl-main)

Copied to clipboard

Challenge: linguistic properties of child-directed speech differ from adult-directed in many ways . linguistic differences between CDS and ADS are retained, but the acoustic properties are similar.
Approach: They compare the task performance of models trained on adult-directed speech and child-directed language . they propose that CDS is optimized for learnability, but not for comprehension .
Outcome: The proposed model trains on adult-directed speech and child-directed language . the model generalizes better on the training register and on synthesized speech .
Multimodal In-context Learning for ASR of Low-resource Languages (2026.findings-acl)

Copied to clipboard

Challenge: In-context learning with large language models addresses this limitation, but prior work focuses on high-resource languages covered during training and text-only settings.
Approach: They propose to use multimodal ICL to learn unseen languages with multimodal learning to improve ASR in large language models.
Outcome: The proposed model outperforms existing models on unseen languages with multimodal ICL (MICL) and cross-lingual transfer learning matches or outperformed models without using target-language data.
Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)

Copied to clipboard

Challenge: Using acoustic data, we develop automatic speech recognition systems for three low resource languages.
Approach: They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework.
Outcome: The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut.
Multilingual Models for ASR in Chibchan Languages (2024.naacl-long)

Copied to clipboard

Challenge: Existing algorithms for low resource-intensive languages are not available for these languages . a paper comparing the performance of different models and algorithms for these extremely low resource languages is presented.
Approach: They propose to fine-tune four ASR algorithms to create monolingual models for Bribri and Cabécar . they then use the best performing algorithm to train joint and transfer learning models for both languages .
Outcome: The proposed algorithms are effective in both Bribri and Cabécar, but especially in Bribri.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations