Challenge: a new study examines how accent information is encoded and propagated in an end-to-end ASR system.
Approach: They propose to use phone probes to analyze phonetic content of representations at each layer.
Outcome: The proposed model is based on a large amount of US-accented English speech and is compared with other models using phone probes.

Similar Papers

Evaluation of Off-the-shelf Speech Recognizers on Different Accents in a Dialogue Domain (2022.lrec-1)

Copied to clipboard

Challenge: Existing automatic speech recognition systems for non-American accents have a much higher error rate than for general american accents.
Approach: They evaluate automatic speech recognition systems on agent-directed speech . they find that the performance is worse for non-American accents than for General American .
Outcome: The ASR systems perform worse for non-American accents than for General American accents . the results suggest that training on non-native English speakers is needed to narrow the performance gap.
Accented Speech Recognition With Accent-specific Codebooks (2023.emnlp-main)

Copied to clipboard

Challenge: Degradation in performance across underrepresented accents is a severe deterrent to inclusive adoption of ASR.
Approach: They propose an approach to adapt speech accents to unseen accents by using cross-attention with a trainable set of codebooks.
Outcome: The proposed approach yields significant performance gains on the seen English accents and unseen accents on the Mozilla Common Voice dataset.
Speech Translation and the End-to-End Promise: Taking Stock of Where We Are (2020.acl-main)

Copied to clipboard

Challenge: Until recently, the only feasible approach to translating acoustic speech signals into text was the cascaded approach.
Approach: They propose a classification of the main challenges of traditional approaches to speech translation . they argue that end-to-end models fall short due to compromises made to address data scarcity .
Outcome: This paper provides a brief survey of the main challenges of traditional approaches in speech translation . it reveals that many end-to-end models fail due to compromises made to address data scarcity.
AccentDB: A Database of Non-Native English Accents to Assist Neural Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: aaron e. sanchez and joe saunders: automatic speech recognition still faces a major challenge . they say accents are a way of pronouncing a language, and speakers always have manner of speaking . esassen: accents can be used to identify non-native speakers of a speech .
Approach: They propose to create a database of speech samples in non-native accents for ASR testing . they also propose to introduce accent neutralization of non- native accents to native accent .
Outcome: The proposed model is compared against human-labelled accent classes and is generalized against human data.
Discovering Canonical Indian English Accents: A Crowdsourcing-based Approach (L18-1)

Copied to clipboard

Challenge: Automated Speech Recognition systems degrade in performance when recognizing accents that are different from the ones in training data.
Approach: They propose to adapt Acoustic Models that are trained on one accent to a target accent by using a small amount of speech data in the target accent.
Outcome: The proposed model can be used to identify accents in Indian English and other languages.
Beyond WER: Probing Whisper’s Sub‐token Decoder Across Diverse Language Resource Levels (2025.emnlp-main)

Copied to clipboard

Challenge: Large multilingual automatic speech recognition models achieve remarkable performance, but the internal mechanisms of the end-to-end pipeline remain underexplored.
Approach: They propose to analyze Whisper's multilingual decoder to uncover systematic decoding disparities masked by aggregate error rates.
Outcome: The proposed model performs better on higher resource languages, but lower resource languages fare worse on these metrics.
Can Edge Probing Tests Reveal Linguistic Knowledge in QA Models? (2022.coling-1)

Copied to clipboard

Challenge: grammatical knowledge is encoded in large pre-trained language models (LMs) this is done through supervised classification tasks to predict the grammamatical properties of a span using only the token representations coming from the LM encoder.
Approach: They propose to use a supervised 'edge probing' task to detect grammatical knowledge in large pre-trained language models (LMs) this is done by encoding grammamatical properties using only token representations coming from the LM encoder.
Outcome: The proposed model performs well when fine-tuned or in adversarial situations where the model is forced to learn wrong correlations.
Advancing African-Accented English Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models (2025.acl-srw)

Copied to clipboard

Challenge: Accents play a pivotal role in shaping human communication, a new study finds . existing ASR systems often perform inadequately, even mispronouncing African names .
Approach: They propose a method that uses epistemic uncertainty to automate annotation to reduce costs and human labor.
Outcome: The proposed method reduces costs and human labor by reducing data annotation and epistemic uncertainty.
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents (2024.findings-eacl)

Copied to clipboard

Challenge: AccentFold uses spatial relationships to improve speech recognition for accented speech . existing methods for accent recognition have been limited due to data scarcity and budget constraints .
Approach: They propose a method that exploits spatial relationships between learned accent embeddings to improve downstream automatic speech recognition.
Outcome: The proposed method outperforms baseline methods in accented speech training.
Evaluating Open-Source ASR Systems: Performance Across Diverse Audio Conditions and Error Correction Methods (2025.coling-main)

Copied to clipboard

Challenge: Automated speech recognition (ASR) systems are able to transcribe spontaneous human conversations with high accuracy.
Approach: They evaluate the accuracy of open source automatic speech recognition systems across conversational speech datasets and explore the potential of ASR ensembling and post-ASR correction methods to improve transcription accuracy.
Outcome: The proposed methods highlight the need for robust error correction techniques and address demographic biases to enhance ASR performance and inclusivity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations