Challenge: Currently, studies on bias and fairness in natural language processing focus on a single language and/or across few attributes (e.g. gender, race). However, biases can manifest differently across languages for individual attributes.
Approach: They adapt existing sentiment bias templates in English to Italian, Chinese, Hebrew, and Spanish for race, religion, nationality, and gender.
Outcome: The proposed model favors groups that are dominant in each language's culture, indicating bias amplification, after multilingual finetuning.

Similar Papers

Social Bias in Multilingual Language Models: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Pretrained multilingual models exhibit the same social bias as models processing English texts.
Approach: They examine the literature on bias evaluation and mitigation approaches in multilingual and non-English contexts and identify gaps in the field.
Outcome: The proposed models perform well on multilingual language understanding benchmarks and are consistent with the current literature.
Cross-lingual Transfer Can Worsen Bias in Sentiment Analysis (2023.emnlp-main)

Copied to clipboard

Challenge: Existing evidence of demographic bias in SA systems is limited to a handful of languages, and it is costly to create supervised training data in a new language.
Approach: They use counterfactual evaluation to test whether gender or racial biases are imported when using cross-lingual transfer . r&r is much more prevalent than gender biase .
Outcome: The proposed model is compared with monolingual systems in five languages and shows that it is biased more than monolingual ones.
Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer (2020.acl-main)

Copied to clipboard

Challenge: Multilingual word embeddings embed words from many languages into a single semantic space such that words with similar meanings are close to each other regardless of the language.
Approach: They propose to use multilingual word embeddings to align embeddable words from multiple languages into a single semantic space so that words with similar meanings are close to each other regardless of the language.
Outcome: The proposed model can be used to learn gender bias in multilingual representations and to improve transfer learning.
Your Stereotypical Mileage May Vary: Practical Challenges of Evaluating Biases in Multiple Languages and Cultural Contexts (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies have identified a gap in the availability of tools and resources to study bias in languages other than English and social contexts outside the north of America.
Approach: They use stereotypes to build a corpus of sentence pairs that cover biases in seven cultural contexts.
Outcome: The proposed resource covers a wide range of languages and cultural settings . it favors sentences that express stereotypes in most bias categories .
Global Voices, Local Biases: Socio-Cultural Prejudices across Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on human biases are heavily skewed towards Western and European languages . despite growing interest in language models, there are several shortcomings in the literature .
Approach: They scale the Word Embedding Association Test to 24 languages and add culturally relevant information for each language.
Outcome: The proposed language models can reflect and often amplify the effects of bias across linguistic, cultural, and societal borders.
Bias and Fairness in Natural Language Processing (D19-2)

Copied to clipboard

Challenge: a tutorial will review the history of bias and fairness studies in machine learning and language processing .
Approach: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it presents recent community effort to quantify and mitigat bias in natural language processing models .
Outcome: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it aims to quantify and mitigate bias in natural language processing models for a wide spectrum of tasks .
Are Pretrained Multilingual Models Equally Fair across Languages? (2022.coling-1)

Copied to clipboard

Challenge: Pretrained multilingual language models can help bridge the digital language divide, enabling high-quality NLP models for lower-resourced languages.
Approach: They propose to use a multilingual dataset to examine whether multilingual models are equally fair across languages.
Outcome: The proposed model enables apples-to-apples comparison across languages of group disparities in multilingual language models.
Speaking Multiple Languages Affects the Moral Bias of Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained multilingual language models are often better on English than other languages . however, they are trained on varying amounts of data for each language .
Approach: They apply the MORALDIRECTION framework to multilingual models and analyse their results . they find that PMLMs encode differing moral biases, but these do not correspond to cultural differences or commonalities in human opinions.
Outcome: The proposed model captures moral norms from English and imposes them on other languages.
Investigating Bias in Multilingual Language Models: Cross-Lingual Transfer of Debiasing Techniques (2023.emnlp-main)

Copied to clipboard

Challenge: Debiasing techniques that target sentence representations are being investigated in multilingual models . a growing interest in addressing bias detection and mitigation in NLP due to their societal implications.
Approach: They examine the transferability of debiasing techniques across different languages within multilingual models by using a dataset from CrowS-Pairs.
Outcome: The proposed techniques reduce bias in English, French, German, and Dutch by 13% . the authors also show that the techniques with additional pretraining exhibit enhanced cross-lingual effectiveness for the languages included in the analyses .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations