Word Error Rate Estimation for Speech Recognition: e-WER (P18-2)

Copied to clipboard

Challenge: Automatic speech recognition (ASR) systems require manual transcription of test data to compute the word error rate (WER).
Approach: They propose an approach to estimate word error rate (e-WER) that does not require a gold-standard transcription of the test set.
Outcome: The proposed approach achieves 16.9% WER root mean squared error across 1,400 sentences.

Similar Papers

Automatic Speech Recognition System-Independent Word Error Rate Estimation (2024.lrec-main)

Copied to clipboard

Challenge: Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition systems.
Approach: They propose a hypothesis generation method for ASR system-dependent WER estimation . they use phonetically similar or linguistically more likely alternative words to generate hypotheses .
Outcome: The proposed method outperforms baseline estimators on in-domain data and out-of-domain on Switchboard and CALLHOME.
WER-BERT: Automatic WER Estimation with BERT in a Balanced Ordinal Classification Paradigm (2021.eacl-main)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are evaluated using Word Error Rate (WER) a higher WER means a lower percentage of errors between the ground truth and the transcription of the system.
Approach: They propose a new balanced paradigm for automatic Word Error Rate estimation using a Librispeech dataset and a Google Cloud's Speech-to-Text API.
Outcome: The proposed approach is more effective than regression in a classification setting, but suffers from heavy class imbalance.
Advocating Character Error Rate for Multilingual ASR Evaluation (2025.findings-naacl)

Copied to clipboard

Challenge: Word error rate (WER) has been used for automatic speech recognition (ASR) evaluations for English datasets for many years.
Approach: They propose to use the character error rate as the primary metric in multilingual ASR evaluation to account for morphologically complex languages.
Outcome: The character error rate (CER) is the primary evaluation metric in multilingual ASR evaluation.
A Benchmark of French ASR Systems Based on Error Severity (2025.coling-main)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) transcription errors are often assessed using metrics that compare them with a reference transcription.
Approach: They propose to categorize transcription errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis.
Outcome: The proposed evaluation categorizes errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis.
On the Robust Approximation of ASR Metrics (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for estimating speech recognition metrics depend on ground truth labels.
Approach: They propose a label-free approach to approximating ASR performance metrics . they embed multimodal embeddings in a unified space for speech and transcription representations .
Outcome: The proposed method outperforms baseline models on speech recognition benchmarks by 50%.
Language-specific Effects on Automatic Speech Recognition Errors for World Englishes (2022.coling-1)

Copied to clipboard

Challenge: Existing systems are not able to meet the needs of speakers of different demographic groups.
Approach: They propose to analyze the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors.
Outcome: The proposed system predicts certain errors from the phonological structure of a speaker’s native language.
CEASR: A Corpus for Evaluating Automatic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are increasingly needed for research and practical applications.
Approach: They propose to use public speech corpora to evaluate the quality of automatic speech recognition (ASR) they calculate an average Word Error Rate (WER) per corpus, per system and per corpor-system pair .
Outcome: The proposed corpus evaluates the quality of automatic speech recognition systems using public speech corpora and transcripts generated by state-of-the-art systems.
WER we are and WER we think we are (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent reports of very low word error rates (WERs) achieved by modern automatic speech recognition systems are skepticism towards the accuracy of modern systems.
Approach: They propose to use a dataset to test automatic speech recognition systems . they propose guidelines for creating real-life datasets with high quality annotations .
Outcome: The proposed system achieves 81% of accuracy on human-chatbot interactions compared to the best reported results on human conversations and public benchmarks.
WER We Stand: Benchmarking Urdu ASR Models (2025.coling-main)

Copied to clipboard

Challenge: This paper analyzes the performance of three ASR models for low-resource languages like Urdu . low-rural languages like urdu have significant gaps in accuracy and reliability .
Approach: They evaluate the performance of three ASR models: Whisper, MMS, and Seamless-M4T . they present the first conversational speech dataset for benchmarking Urdu ASR systems .
Outcome: The proposed model families outperform Whisper, MMS, and Seamless-M4T on two types of speech datasets.
Lost in Transcription: Identifying and Quantifying the Accuracy Biases of Automatic Speech Recognition Systems Against Disfluent Speech (2024.naacl-long)

Copied to clipboard

Challenge: Automatic speech recognition systems fail to accurately interpret speech patterns deviating from typical fluency, leading to critical usability issues and misinterpretations.
Approach: They evaluate six leading automatic speech recognition systems based on a real-world dataset and a synthetic dataset derived from the widely-used LibriSpeech benchmark.
Outcome: The six leading speech recognition systems were evaluated on a real-world dataset and a synthetic dataset derived from the widely-used LibriSpeech benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations