Challenge: Automatic Speech Recognition (ASR) transcription errors are often assessed using metrics that compare them with a reference transcription.
Approach: They propose to categorize transcription errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis.
Outcome: The proposed evaluation categorizes errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis.

Similar Papers

Advocating Character Error Rate for Multilingual ASR Evaluation (2025.findings-naacl)

Copied to clipboard

Challenge: Word error rate (WER) has been used for automatic speech recognition (ASR) evaluations for English datasets for many years.
Approach: They propose to use the character error rate as the primary metric in multilingual ASR evaluation to account for morphologically complex languages.
Outcome: The character error rate (CER) is the primary evaluation metric in multilingual ASR evaluation.
Word Error Rate Estimation for Speech Recognition: e-WER (P18-2)

Copied to clipboard

Challenge: Automatic speech recognition (ASR) systems require manual transcription of test data to compute the word error rate (WER).
Approach: They propose an approach to estimate word error rate (e-WER) that does not require a gold-standard transcription of the test set.
Outcome: The proposed approach achieves 16.9% WER root mean squared error across 1,400 sentences.
WER-BERT: Automatic WER Estimation with BERT in a Balanced Ordinal Classification Paradigm (2021.eacl-main)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are evaluated using Word Error Rate (WER) a higher WER means a lower percentage of errors between the ground truth and the transcription of the system.
Approach: They propose a new balanced paradigm for automatic Word Error Rate estimation using a Librispeech dataset and a Google Cloud's Speech-to-Text API.
Outcome: The proposed approach is more effective than regression in a classification setting, but suffers from heavy class imbalance.
Language-specific Effects on Automatic Speech Recognition Errors for World Englishes (2022.coling-1)

Copied to clipboard

Challenge: Existing systems are not able to meet the needs of speakers of different demographic groups.
Approach: They propose to analyze the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors.
Outcome: The proposed system predicts certain errors from the phonological structure of a speaker’s native language.
CEASR: A Corpus for Evaluating Automatic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are increasingly needed for research and practical applications.
Approach: They propose to use public speech corpora to evaluate the quality of automatic speech recognition (ASR) they calculate an average Word Error Rate (WER) per corpus, per system and per corpor-system pair .
Outcome: The proposed corpus evaluates the quality of automatic speech recognition systems using public speech corpora and transcripts generated by state-of-the-art systems.
Evaluating Open-Source ASR Systems: Performance Across Diverse Audio Conditions and Error Correction Methods (2025.coling-main)

Copied to clipboard

Challenge: Automated speech recognition (ASR) systems are able to transcribe spontaneous human conversations with high accuracy.
Approach: They evaluate the accuracy of open source automatic speech recognition systems across conversational speech datasets and explore the potential of ASR ensembling and post-ASR correction methods to improve transcription accuracy.
Outcome: The proposed methods highlight the need for robust error correction techniques and address demographic biases to enhance ASR performance and inclusivity.
WER we are and WER we think we are (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent reports of very low word error rates (WERs) achieved by modern automatic speech recognition systems are skepticism towards the accuracy of modern systems.
Approach: They propose to use a dataset to test automatic speech recognition systems . they propose guidelines for creating real-life datasets with high quality annotations .
Outcome: The proposed system achieves 81% of accuracy on human-chatbot interactions compared to the best reported results on human conversations and public benchmarks.
That doesn’t sound right: Evaluating speech transcription quality in field linguistics corpora (2025.acl-short)

Copied to clipboard

Challenge: Automated speech recognition (ASR) is a popular tool for documenting languages, but field linguists do not have the data to train robust models.
Approach: They propose to use fieldwork data to identify speech transcriptions that may be unsuitable for training ASR models.
Outcome: The proposed measures can be used to identify transcriptions with characteristics common in field data but could be detrimental to ASR training.
WER We Stand: Benchmarking Urdu ASR Models (2025.coling-main)

Copied to clipboard

Challenge: This paper analyzes the performance of three ASR models for low-resource languages like Urdu . low-rural languages like urdu have significant gaps in accuracy and reliability .
Approach: They evaluate the performance of three ASR models: Whisper, MMS, and Seamless-M4T . they present the first conversational speech dataset for benchmarking Urdu ASR systems .
Outcome: The proposed model families outperform Whisper, MMS, and Seamless-M4T on two types of speech datasets.
Automatic Speech Recognition System-Independent Word Error Rate Estimation (2024.lrec-main)

Copied to clipboard

Challenge: Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition systems.
Approach: They propose a hypothesis generation method for ASR system-dependent WER estimation . they use phonetically similar or linguistically more likely alternative words to generate hypotheses .
Outcome: The proposed method outperforms baseline estimators on in-domain data and out-of-domain on Switchboard and CALLHOME.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations