Challenge: State-of-the-art automatic speech recognition systems exhibit disparate performance on varying speech accents.
Approach: They propose to use submodular mutual information to find the most informative set of utterances matching a target accent within a fixed budget.
Outcome: The proposed model is 3-5 times more label-efficient on the Indic-TTS and L2 datasets than other methods.

Similar Papers

On Mitigating Performance Disparities in Multilingual Speech Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: Automatic Speech Recognition systems are not always equally effective for all users, and gender disparity in their performance is a significant concern.
Approach: They compare performance of different fine-tuning algorithms for multilingual speech recognition across languages and genders.
Outcome: The proposed algorithms improve performance and parity across languages and languages.
Improving Language Identification for Code-Switched Speech: The Pivotal Role of Accented English (2026.findings-eacl)

Copied to clipboard

Challenge: Existing models fail to identify English spoken with the accent of the matrix (dominant) language.
Approach: They propose to fine tune existing LID models with accented English to improve code-switched LID . they use a metric that captures relative ranking of identified languages often overlooked by traditional metrics.
Outcome: The proposed model can be fine tuned with small amounts of accented English without degrading performance on monolingual speech.
Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation (2023.eacl-main)

Copied to clipboard

Challenge: Automatic speech recognition data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set.
Approach: They propose to use hold-speaker(s)-out partitioning to partition data for five languages . utterance duration and intensity are more predictive factors of variability .
Outcome: The proposed method can produce results that do not reflect model performance on unseen data or speakers.
Evaluation of Off-the-shelf Speech Recognizers on Different Accents in a Dialogue Domain (2022.lrec-1)

Copied to clipboard

Challenge: Existing automatic speech recognition systems for non-American accents have a much higher error rate than for general american accents.
Approach: They evaluate automatic speech recognition systems on agent-directed speech . they find that the performance is worse for non-American accents than for General American .
Outcome: The ASR systems perform worse for non-American accents than for General American accents . the results suggest that training on non-native English speakers is needed to narrow the performance gap.
Advancing African-Accented English Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models (2025.acl-srw)

Copied to clipboard

Challenge: Accents play a pivotal role in shaping human communication, a new study finds . existing ASR systems often perform inadequately, even mispronouncing African names .
Approach: They propose a method that uses epistemic uncertainty to automate annotation to reduce costs and human labor.
Outcome: The proposed method reduces costs and human labor by reducing data annotation and epistemic uncertainty.
Fairness in Automatic Speech Recognition Isn’t a One-Size-Fits-All (2025.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained speech models like Whisper exhibit inconsistent group-level performance that varies across domains.
Approach: They fine-tune a Whisper model on the Fair-Speech corpus using basic fine- tuning, demographic rebalancing, gender-swapped data augmentation and a novel contrastive learning objective.
Outcome: The proposed method achieves stable, cross-domain fairness improvements without changes to the training data distribution and with minimal accuracy trade-offs.
Partitioned Gradient Matching-based Data Subset Selection for Compute-Efficient Robust ASR Training (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing DSS algorithms for RNN-T have a high cost and performance degradation.
Approach: They propose a distributable DSS algorithm for RNN-T that can be used to train a subset of training data.
Outcome: The proposed algorithm achieves between 3x to 6x speedup with only a small accuracy degradation even in settings where the training data is corrupted with noise.
Dialectal Coverage And Generalization in Arabic Speech Recognition (2025.acl-long)

Copied to clipboard

Challenge: Existing ASR systems cover the modern standard Arabic variety but fail to cover the multitude of spoken variants.
Approach: They propose a suite of automatic speech recognition models optimized to recognize multiple variants of spoken Arabic.
Outcome: The proposed models show coverage and performance gains compared to prior models.
AccentDB: A Database of Non-Native English Accents to Assist Neural Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: aaron e. sanchez and joe saunders: automatic speech recognition still faces a major challenge . they say accents are a way of pronouncing a language, and speakers always have manner of speaking . esassen: accents can be used to identify non-native speakers of a speech .
Approach: They propose to create a database of speech samples in non-native accents for ASR testing . they also propose to introduce accent neutralization of non- native accents to native accent .
Outcome: The proposed model is compared against human-labelled accent classes and is generalized against human data.
Evaluation of Off-the-shelf Speech Recognizers Across Diverse Dialogue Domains (2020.lrec-1)

Copied to clipboard

Challenge: a recent study evaluated off-the-shelf automatic speech recognition systems . current state-of-the art systems perform poorly in domains that require special vocabulary and language models .
Approach: They evaluate off-the-shelf automatic speech recognition systems across different dialogue domains . they use data collected from deployed spoken dialogue systems and human-human conversations .
Outcome: The evaluation is aimed at non-experts with limited experience in speech recognition . the results show that the performance of each speech recognizer can vary significantly depending on the domain .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations