DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation (2023.acl-long)
Copied to clipboard
Suraj Kothawade, Anmol Mekala, D.Chandra Sekhara Hetha Havya, Mayank Kothyari, Rishabh Iyer, Ganesh Ramakrishnan, Preethi Jyothi
| Challenge: | State-of-the-art automatic speech recognition systems exhibit disparate performance on varying speech accents. |
| Approach: | They propose to use submodular mutual information to find the most informative set of utterances matching a target accent within a fixed budget. |
| Outcome: | The proposed model is 3-5 times more label-efficient on the Indic-TTS and L2 datasets than other methods. |
Similar Papers
On Mitigating Performance Disparities in Multilingual Speech Recognition (2024.emnlp-main)
Copied to clipboard
| Challenge: | Automatic Speech Recognition systems are not always equally effective for all users, and gender disparity in their performance is a significant concern. |
| Approach: | They compare performance of different fine-tuning algorithms for multilingual speech recognition across languages and genders. |
| Outcome: | The proposed algorithms improve performance and parity across languages and languages. |
Improving Language Identification for Code-Switched Speech: The Pivotal Role of Accented English (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing models fail to identify English spoken with the accent of the matrix (dominant) language. |
| Approach: | They propose to fine tune existing LID models with accented English to improve code-switched LID . they use a metric that captures relative ranking of identified languages often overlooked by traditional metrics. |
| Outcome: | The proposed model can be fine tuned with small amounts of accented English without degrading performance on monolingual speech. |
Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation (2023.eacl-main)
Copied to clipboard
| Challenge: | Automatic speech recognition data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set. |
| Approach: | They propose to use hold-speaker(s)-out partitioning to partition data for five languages . utterance duration and intensity are more predictive factors of variability . |
| Outcome: | The proposed method can produce results that do not reflect model performance on unseen data or speakers. |
Evaluation of Off-the-shelf Speech Recognizers on Different Accents in a Dialogue Domain (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing automatic speech recognition systems for non-American accents have a much higher error rate than for general american accents. |
| Approach: | They evaluate automatic speech recognition systems on agent-directed speech . they find that the performance is worse for non-American accents than for General American . |
| Outcome: | The ASR systems perform worse for non-American accents than for General American accents . the results suggest that training on non-native English speakers is needed to narrow the performance gap. |
Advancing African-Accented English Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models (2025.acl-srw)
Copied to clipboard
| Challenge: | Accents play a pivotal role in shaping human communication, a new study finds . existing ASR systems often perform inadequately, even mispronouncing African names . |
| Approach: | They propose a method that uses epistemic uncertainty to automate annotation to reduce costs and human labor. |
| Outcome: | The proposed method reduces costs and human labor by reducing data annotation and epistemic uncertainty. |
Fairness in Automatic Speech Recognition Isn’t a One-Size-Fits-All (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Pre-trained speech models like Whisper exhibit inconsistent group-level performance that varies across domains. |
| Approach: | They fine-tune a Whisper model on the Fair-Speech corpus using basic fine- tuning, demographic rebalancing, gender-swapped data augmentation and a novel contrastive learning objective. |
| Outcome: | The proposed method achieves stable, cross-domain fairness improvements without changes to the training data distribution and with minimal accuracy trade-offs. |
Partitioned Gradient Matching-based Data Subset Selection for Compute-Efficient Robust ASR Training (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing DSS algorithms for RNN-T have a high cost and performance degradation. |
| Approach: | They propose a distributable DSS algorithm for RNN-T that can be used to train a subset of training data. |
| Outcome: | The proposed algorithm achieves between 3x to 6x speedup with only a small accuracy degradation even in settings where the training data is corrupted with noise. |
Dialectal Coverage And Generalization in Arabic Speech Recognition (2025.acl-long)
Copied to clipboard
| Challenge: | Existing ASR systems cover the modern standard Arabic variety but fail to cover the multitude of spoken variants. |
| Approach: | They propose a suite of automatic speech recognition models optimized to recognize multiple variants of spoken Arabic. |
| Outcome: | The proposed models show coverage and performance gains compared to prior models. |
AccentDB: A Database of Non-Native English Accents to Assist Neural Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | aaron e. sanchez and joe saunders: automatic speech recognition still faces a major challenge . they say accents are a way of pronouncing a language, and speakers always have manner of speaking . esassen: accents can be used to identify non-native speakers of a speech . |
| Approach: | They propose to create a database of speech samples in non-native accents for ASR testing . they also propose to introduce accent neutralization of non- native accents to native accent . |
| Outcome: | The proposed model is compared against human-labelled accent classes and is generalized against human data. |
Evaluation of Off-the-shelf Speech Recognizers Across Diverse Dialogue Domains (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study evaluated off-the-shelf automatic speech recognition systems . current state-of-the art systems perform poorly in domains that require special vocabulary and language models . |
| Approach: | They evaluate off-the-shelf automatic speech recognition systems across different dialogue domains . they use data collected from deployed spoken dialogue systems and human-human conversations . |
| Outcome: | The evaluation is aimed at non-experts with limited experience in speech recognition . the results show that the performance of each speech recognizer can vary significantly depending on the domain . |