How Accents Confound: Probing for Accent Information in End-to-End Speech Recognition Systems (2020.acl-main)
Copied to clipboard
| Challenge: | a new study examines how accent information is encoded and propagated in an end-to-end ASR system. |
| Approach: | They propose to use phone probes to analyze phonetic content of representations at each layer. |
| Outcome: | The proposed model is based on a large amount of US-accented English speech and is compared with other models using phone probes. |
Similar Papers
Evaluation of Off-the-shelf Speech Recognizers on Different Accents in a Dialogue Domain (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing automatic speech recognition systems for non-American accents have a much higher error rate than for general american accents. |
| Approach: | They evaluate automatic speech recognition systems on agent-directed speech . they find that the performance is worse for non-American accents than for General American . |
| Outcome: | The ASR systems perform worse for non-American accents than for General American accents . the results suggest that training on non-native English speakers is needed to narrow the performance gap. |
Accented Speech Recognition With Accent-specific Codebooks (2023.emnlp-main)
Copied to clipboard
| Challenge: | Degradation in performance across underrepresented accents is a severe deterrent to inclusive adoption of ASR. |
| Approach: | They propose an approach to adapt speech accents to unseen accents by using cross-attention with a trainable set of codebooks. |
| Outcome: | The proposed approach yields significant performance gains on the seen English accents and unseen accents on the Mozilla Common Voice dataset. |
Speech Translation and the End-to-End Promise: Taking Stock of Where We Are (2020.acl-main)
Copied to clipboard
| Challenge: | Until recently, the only feasible approach to translating acoustic speech signals into text was the cascaded approach. |
| Approach: | They propose a classification of the main challenges of traditional approaches to speech translation . they argue that end-to-end models fall short due to compromises made to address data scarcity . |
| Outcome: | This paper provides a brief survey of the main challenges of traditional approaches in speech translation . it reveals that many end-to-end models fail due to compromises made to address data scarcity. |
AccentDB: A Database of Non-Native English Accents to Assist Neural Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | aaron e. sanchez and joe saunders: automatic speech recognition still faces a major challenge . they say accents are a way of pronouncing a language, and speakers always have manner of speaking . esassen: accents can be used to identify non-native speakers of a speech . |
| Approach: | They propose to create a database of speech samples in non-native accents for ASR testing . they also propose to introduce accent neutralization of non- native accents to native accent . |
| Outcome: | The proposed model is compared against human-labelled accent classes and is generalized against human data. |
Discovering Canonical Indian English Accents: A Crowdsourcing-based Approach (L18-1)
Copied to clipboard
| Challenge: | Automated Speech Recognition systems degrade in performance when recognizing accents that are different from the ones in training data. |
| Approach: | They propose to adapt Acoustic Models that are trained on one accent to a target accent by using a small amount of speech data in the target accent. |
| Outcome: | The proposed model can be used to identify accents in Indian English and other languages. |
Beyond WER: Probing Whisper’s Sub‐token Decoder Across Diverse Language Resource Levels (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large multilingual automatic speech recognition models achieve remarkable performance, but the internal mechanisms of the end-to-end pipeline remain underexplored. |
| Approach: | They propose to analyze Whisper's multilingual decoder to uncover systematic decoding disparities masked by aggregate error rates. |
| Outcome: | The proposed model performs better on higher resource languages, but lower resource languages fare worse on these metrics. |
Can Edge Probing Tests Reveal Linguistic Knowledge in QA Models? (2022.coling-1)
Copied to clipboard
| Challenge: | grammatical knowledge is encoded in large pre-trained language models (LMs) this is done through supervised classification tasks to predict the grammamatical properties of a span using only the token representations coming from the LM encoder. |
| Approach: | They propose to use a supervised 'edge probing' task to detect grammatical knowledge in large pre-trained language models (LMs) this is done by encoding grammamatical properties using only token representations coming from the LM encoder. |
| Outcome: | The proposed model performs well when fine-tuned or in adversarial situations where the model is forced to learn wrong correlations. |
Advancing African-Accented English Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models (2025.acl-srw)
Copied to clipboard
| Challenge: | Accents play a pivotal role in shaping human communication, a new study finds . existing ASR systems often perform inadequately, even mispronouncing African names . |
| Approach: | They propose a method that uses epistemic uncertainty to automate annotation to reduce costs and human labor. |
| Outcome: | The proposed method reduces costs and human labor by reducing data annotation and epistemic uncertainty. |
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents (2024.findings-eacl)
Copied to clipboard
| Challenge: | AccentFold uses spatial relationships to improve speech recognition for accented speech . existing methods for accent recognition have been limited due to data scarcity and budget constraints . |
| Approach: | They propose a method that exploits spatial relationships between learned accent embeddings to improve downstream automatic speech recognition. |
| Outcome: | The proposed method outperforms baseline methods in accented speech training. |
Evaluating Open-Source ASR Systems: Performance Across Diverse Audio Conditions and Error Correction Methods (2025.coling-main)
Copied to clipboard
| Challenge: | Automated speech recognition (ASR) systems are able to transcribe spontaneous human conversations with high accuracy. |
| Approach: | They evaluate the accuracy of open source automatic speech recognition systems across conversational speech datasets and explore the potential of ASR ensembling and post-ASR correction methods to improve transcription accuracy. |
| Outcome: | The proposed methods highlight the need for robust error correction techniques and address demographic biases to enhance ASR performance and inclusivity. |