Challenge: Existing systems are not able to meet the needs of speakers of different demographic groups.
Approach: They propose to analyze the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors.
Outcome: The proposed system predicts certain errors from the phonological structure of a speaker’s native language.

Similar Papers

Advocating Character Error Rate for Multilingual ASR Evaluation (2025.findings-naacl)

Copied to clipboard

Challenge: Word error rate (WER) has been used for automatic speech recognition (ASR) evaluations for English datasets for many years.
Approach: They propose to use the character error rate as the primary metric in multilingual ASR evaluation to account for morphologically complex languages.
Outcome: The character error rate (CER) is the primary evaluation metric in multilingual ASR evaluation.
Evaluating Open-Source ASR Systems: Performance Across Diverse Audio Conditions and Error Correction Methods (2025.coling-main)

Copied to clipboard

Challenge: Automated speech recognition (ASR) systems are able to transcribe spontaneous human conversations with high accuracy.
Approach: They evaluate the accuracy of open source automatic speech recognition systems across conversational speech datasets and explore the potential of ASR ensembling and post-ASR correction methods to improve transcription accuracy.
Outcome: The proposed methods highlight the need for robust error correction techniques and address demographic biases to enhance ASR performance and inclusivity.
Evaluation of Off-the-shelf Speech Recognizers on Different Accents in a Dialogue Domain (2022.lrec-1)

Copied to clipboard

Challenge: Existing automatic speech recognition systems for non-American accents have a much higher error rate than for general american accents.
Approach: They evaluate automatic speech recognition systems on agent-directed speech . they find that the performance is worse for non-American accents than for General American .
Outcome: The ASR systems perform worse for non-American accents than for General American accents . the results suggest that training on non-native English speakers is needed to narrow the performance gap.
Modeling Gender and Dialect Bias in Automatic Speech Recognition (2024.findings-emnlp)

Copied to clipboard

Challenge: Dialect and gender-based biases have become an area of concern in language-dependent AI systems.
Approach: They construct a podcast audio dataset and evaluate its performance . they then refine the models to better understand how finetuning can impact performance.
Outcome: The proposed model improves on 13 hours of podcast audio transcribed by speakers of four US-based English dialects.
Language technology practitioners as language managers: arbitrating data bias and predictive bias in ASR (2022.lrec-1)

Copied to clipboard

Challenge: despite natural language variation, automatic speech recognition systems perform worse on non-standardised and marginalised language varieties.
Approach: They propose a re-framing of language resources as (public) infrastructure for speech communities . authors propose rethinking of algorithms to address the origins and harms of bias .
Outcome: The proposed approach aims to understand the origins and harms of algorithmic bias and how it can be mitigated.
Why Aren’t We NER Yet? Artifacts of ASR Errors in Named Entity Recognition in Spontaneous Speech Transcripts (2023.acl-long)

Copied to clipboard

Challenge: despite advances in language models, the transcript of spontaneous human-human conversations remains an insurmountable challenge for most models.
Approach: They examine the relationship between ASR and NER errors which limit NER models' ability to recover entity mentions from spontaneous speech transcripts.
Outcome: The proposed model fails even if no word errors are introduced by the ASR . the proposed model's performance deteriorates when applied to the ASL outputs .
The Influence of Automatic Speech Recognition on Linguistic Features and Automatic Alzheimer’s Disease Detection from Spontaneous Speech (2024.lrec-main)

Copied to clipboard

Challenge: Existing biomarkers for AD diagnosis can only be applied to relatively small sample sizes due to limited availability, excessive costs and invasive nature.
Approach: They compare automatic speech recognition systems in terms of Word Error Rate (WER) using a publicly available benchmark dataset of speech recordings of AD patients and controls.
Outcome: The proposed method improves classification performance by replacing manual transcriptions with ASR output.
A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models.
Approach: They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages.
Outcome: The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data.
Lost in Transcription: Identifying and Quantifying the Accuracy Biases of Automatic Speech Recognition Systems Against Disfluent Speech (2024.naacl-long)

Copied to clipboard

Challenge: Automatic speech recognition systems fail to accurately interpret speech patterns deviating from typical fluency, leading to critical usability issues and misinterpretations.
Approach: They evaluate six leading automatic speech recognition systems based on a real-world dataset and a synthetic dataset derived from the widely-used LibriSpeech benchmark.
Outcome: The six leading speech recognition systems were evaluated on a real-world dataset and a synthetic dataset derived from the widely-used LibriSpeech benchmark.
What is lost in Normalization? Exploring Pitfalls in Multilingual ASR Model Evaluations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing text normalization routines that target Indic scripts are flawed when applied to multilingual automatic speech recognition models.
Approach: They propose to develop text normalization routines that leverage native linguistic expertise to ensure more robust and accurate evaluations of multilingual automatic speech recognition models.
Outcome: The proposed normalization routines can be leveraged to improve performance metrics for Indic languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations