Advocating Character Error Rate for Multilingual ASR Evaluation (2025.findings-naacl)
Copied to clipboard
| Challenge: | Word error rate (WER) has been used for automatic speech recognition (ASR) evaluations for English datasets for many years. |
| Approach: | They propose to use the character error rate as the primary metric in multilingual ASR evaluation to account for morphologically complex languages. |
| Outcome: | The character error rate (CER) is the primary evaluation metric in multilingual ASR evaluation. |
Similar Papers
A Benchmark of French ASR Systems Based on Error Severity (2025.coling-main)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) transcription errors are often assessed using metrics that compare them with a reference transcription. |
| Approach: | They propose to categorize transcription errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis. |
| Outcome: | The proposed evaluation categorizes errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis. |
Language-specific Effects on Automatic Speech Recognition Errors for World Englishes (2022.coling-1)
Copied to clipboard
| Challenge: | Existing systems are not able to meet the needs of speakers of different demographic groups. |
| Approach: | They propose to analyze the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors. |
| Outcome: | The proposed system predicts certain errors from the phonological structure of a speaker’s native language. |
Word Error Rate Estimation for Speech Recognition: e-WER (P18-2)
Copied to clipboard
| Challenge: | Automatic speech recognition (ASR) systems require manual transcription of test data to compute the word error rate (WER). |
| Approach: | They propose an approach to estimate word error rate (e-WER) that does not require a gold-standard transcription of the test set. |
| Outcome: | The proposed approach achieves 16.9% WER root mean squared error across 1,400 sentences. |
On the Robust Approximation of ASR Metrics (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for estimating speech recognition metrics depend on ground truth labels. |
| Approach: | They propose a label-free approach to approximating ASR performance metrics . they embed multimodal embeddings in a unified space for speech and transcription representations . |
| Outcome: | The proposed method outperforms baseline models on speech recognition benchmarks by 50%. |
CEASR: A Corpus for Evaluating Automatic Speech Recognition (2020.lrec-1)
Copied to clipboard
Malgorzata Anna Ulasik, Manuela Hürlimann, Fabian Germann, Esin Gedik, Fernando Benites, Mark Cieliebak
| Challenge: | Automatic Speech Recognition (ASR) systems are increasingly needed for research and practical applications. |
| Approach: | They propose to use public speech corpora to evaluate the quality of automatic speech recognition (ASR) they calculate an average Word Error Rate (WER) per corpus, per system and per corpor-system pair . |
| Outcome: | The proposed corpus evaluates the quality of automatic speech recognition systems using public speech corpora and transcripts generated by state-of-the-art systems. |
Automatic Speech Recognition System-Independent Word Error Rate Estimation (2024.lrec-main)
Copied to clipboard
| Challenge: | Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition systems. |
| Approach: | They propose a hypothesis generation method for ASR system-dependent WER estimation . they use phonetically similar or linguistically more likely alternative words to generate hypotheses . |
| Outcome: | The proposed method outperforms baseline estimators on in-domain data and out-of-domain on Switchboard and CALLHOME. |
WER we are and WER we think we are (2020.findings-emnlp)
Copied to clipboard
Piotr Szymański, Piotr Żelasko, Mikolaj Morzy, Adrian Szymczak, Marzena Żyła-Hoppe, Joanna Banaszczak, Lukasz Augustyniak, Jan Mizgajski, Yishay Carmiel
| Challenge: | Recent reports of very low word error rates (WERs) achieved by modern automatic speech recognition systems are skepticism towards the accuracy of modern systems. |
| Approach: | They propose to use a dataset to test automatic speech recognition systems . they propose guidelines for creating real-life datasets with high quality annotations . |
| Outcome: | The proposed system achieves 81% of accuracy on human-chatbot interactions compared to the best reported results on human conversations and public benchmarks. |
What is lost in Normalization? Exploring Pitfalls in Multilingual ASR Model Evaluations (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing text normalization routines that target Indic scripts are flawed when applied to multilingual automatic speech recognition models. |
| Approach: | They propose to develop text normalization routines that leverage native linguistic expertise to ensure more robust and accurate evaluations of multilingual automatic speech recognition models. |
| Outcome: | The proposed normalization routines can be leveraged to improve performance metrics for Indic languages. |
A Human Evaluation of AMR-to-English Generation Systems (2020.coling-main)
Copied to clipboard
| Challenge: | a recent human evaluation of AMR generation systems is compared to automated metrics. |
| Approach: | They propose a human evaluation which collects fluency and adequacy scores and categorization of error types for AMR generation systems. |
| Outcome: | The results show that human evaluations are more nuanced than automated metrics. |
WER We Stand: Benchmarking Urdu ASR Models (2025.coling-main)
Copied to clipboard
| Challenge: | This paper analyzes the performance of three ASR models for low-resource languages like Urdu . low-rural languages like urdu have significant gaps in accuracy and reliability . |
| Approach: | They evaluate the performance of three ASR models: Whisper, MMS, and Seamless-M4T . they present the first conversational speech dataset for benchmarking Urdu ASR systems . |
| Outcome: | The proposed model families outperform Whisper, MMS, and Seamless-M4T on two types of speech datasets. |