Challenge: a new class of multitasks, multilingual neural networks, has recently pushed the boundaries of speech-related tasks.
Approach: They evaluate performance of two widely used multilingual automatic speech recognition models . they find clear gender disparities, with the advantaged group varying across languages .
Outcome: The proposed models are compared on 19 languages from eight language families and two speaking conditions.

Similar Papers

Modeling Gender and Dialect Bias in Automatic Speech Recognition (2024.findings-emnlp)

Copied to clipboard

Challenge: Dialect and gender-based biases have become an area of concern in language-dependent AI systems.
Approach: They construct a podcast audio dataset and evaluate its performance . they then refine the models to better understand how finetuning can impact performance.
Outcome: The proposed model improves on 13 hours of podcast audio transcribed by speakers of four US-based English dialects.
On Mitigating Performance Disparities in Multilingual Speech Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: Automatic Speech Recognition systems are not always equally effective for all users, and gender disparity in their performance is a significant concern.
Approach: They compare performance of different fine-tuning algorithms for multilingual speech recognition across languages and genders.
Outcome: The proposed algorithms improve performance and parity across languages and languages.
A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models.
Approach: They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages.
Outcome: The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data.
Different Speech Translation Models Encode and Translate Speaker Gender Differently (2025.acl-short)

Copied to clipboard

Challenge: Recent studies on interpreting the hidden states of speech models have shown their ability to capture speaker-specific features, including gender.
Approach: They propose to use probing methods to assess gender encoding across ST models.
Outcome: The proposed models capture speaker-specific features, including gender, while older models do not . low gender encoding capabilities result in systems’ tendency toward a masculine default, a translation bias that is more pronounced in newer architectures.
Group Fairness in Multilingual Speech Recognition Models (2024.findings-naacl)

Copied to clipboard

Challenge: a new study evaluates the performance disparities of ASR models across languages and demographics . a large amount of data is required to mitigate performance disparity, but this is computationally expensive .
Approach: They evaluate the performance disparity of ASR models using a multilingual dataset . they find that model size correlates logarithmically with worst-case performance disparities .
Outcome: The proposed models exhibit significant performance disparities across binary genders for adolescents.
Language-specific Effects on Automatic Speech Recognition Errors for World Englishes (2022.coling-1)

Copied to clipboard

Challenge: Existing systems are not able to meet the needs of speakers of different demographic groups.
Approach: They propose to analyze the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors.
Outcome: The proposed system predicts certain errors from the phonological structure of a speaker’s native language.
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for measuring gender stereotypical bias in language models are inconsistencies . lack of explicit standards in data gathering can have detrimental effects on results .
Approach: They propose that currently available benchmarks capture only partial facets of gender stereotypes . they apply a framework from social psychology to balance data across components of gender stereotypes based on stereotypical benchmarks.
Outcome: The proposed framework improves correlation between different benchmarks by using simple balancing techniques.
Fairness in Automatic Speech Recognition Isn’t a One-Size-Fits-All (2025.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained speech models like Whisper exhibit inconsistent group-level performance that varies across domains.
Approach: They fine-tune a Whisper model on the Fair-Speech corpus using basic fine- tuning, demographic rebalancing, gender-swapped data augmentation and a novel contrastive learning objective.
Outcome: The proposed method achieves stable, cross-domain fairness improvements without changes to the training data distribution and with minimal accuracy trade-offs.
Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus (2020.acl-main)

Copied to clipboard

Challenge: a growing number of studies have examined the issue of gender bias in speech translation . a gender bias is a systemic problem that reproduces gender stereotypes discriminating women.
Approach: They present the first thorough investigation of gender bias in speech translation . they compare audio technologies for English-Italian/French translations .
Outcome: The proposed method compares different technologies on two languages, English and French.
On Evaluating and Mitigating Gender Biases in Multilingual Settings (2023.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks and resources for evaluating gender biases in multilingual settings are limited.
Approach: They propose to extend DisCo to different Indian languages using human annotations to evaluate gender biases in multilingual models.
Outcome: The proposed benchmarks and mitigation techniques are extended beyond English to evaluate gender biases in multilingual models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations