Challenge: We introduce 1,679 sentence pairs in French that cover stereotypes in ten types of bias like gender and age.
Approach: They build on the US-centered CrowS-pairs dataset to create a multilingual stereotypes dataset that allows for comparability across languages and cultures.
Outcome: The proposed dataset allows for comparability across languages while characterizing biases that are specific to each country and language.

Similar Papers

CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models use cultural biases implicitly, causing harm . identifying and quantifying learnt biase enables us to measure progress .
Approach: They propose a benchmark to measure social bias in pretrained language models . they use 1508 examples that cover stereotypes dealing with nine types of bias .
Outcome: The proposed benchmark focuses on stereotypes about historically disadvantaged groups and contrasts them with advantaged groups.
Your Stereotypical Mileage May Vary: Practical Challenges of Evaluating Biases in Multiple Languages and Cultural Contexts (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies have identified a gap in the availability of tools and resources to study bias in languages other than English and social contexts outside the north of America.
Approach: They use stereotypes to build a corpus of sentence pairs that cover biases in seven cultural contexts.
Outcome: The proposed resource covers a wide range of languages and cultural settings . it favors sentences that express stereotypes in most bias categories .
An Information-Theoretic Approach and Dataset for Probing Gender Stereotypes in Multilingual Masked Language Models (2022.findings-naacl)

Copied to clipboard

Challenge: Pretrained language models (PLMs) have been shown to encapsulate social biases, including those relating to gender and race.
Approach: They propose a new bias measure based on Jensen–Shannon divergence that retains more information from the model output probabilities than other previously proposed bias measures.
Outcome: The proposed measure outperforms CrowS-Pairs and other similar measures for non-English datasets.
Social Bias in Multilingual Language Models: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Pretrained multilingual models exhibit the same social bias as models processing English texts.
Approach: They examine the literature on bias evaluation and mitigation approaches in multilingual and non-English contexts and identify gaps in the field.
Outcome: The proposed models perform well on multilingual language understanding benchmarks and are consistent with the current literature.
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing research on stereotypes in large language models is limited and focuses on African Ameri- F.
Approach: They propose to use global bias to probe a set of large language models via perplexity to determine how certain stereotypes are represented in the model's internal representations.
Outcome: The proposed model amplifys harmful stereotypes and shows that the demographic groups associated with stereotypes remain consistent across model likelihoods and outputs.
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for measuring gender stereotypical bias in language models are inconsistencies . lack of explicit standards in data gathering can have detrimental effects on results .
Approach: They propose that currently available benchmarks capture only partial facets of gender stereotypes . they apply a framework from social psychology to balance data across components of gender stereotypes based on stereotypical benchmarks.
Outcome: The proposed framework improves correlation between different benchmarks by using simple balancing techniques.
In-Depth Look at Word Filling Societal Bias Measures (2023.eacl-main)

Copied to clipboard

Challenge: Language models (LMs) are ubiquitous in current NLP and have brought undeniable performance improvements for many tasks.
Approach: They propose to use word filling prompts to evaluate language models' behavior to find out if they are valid.
Outcome: The proposed measures produce unexpected and illogical results when appropriate control group samples are constructed.
Comparing Biases and the Impact of Multilingual Training across Multiple Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, studies on bias and fairness in natural language processing focus on a single language and/or across few attributes (e.g. gender, race). However, biases can manifest differently across languages for individual attributes.
Approach: They adapt existing sentiment bias templates in English to Italian, Chinese, Hebrew, and Spanish for race, religion, nationality, and gender.
Outcome: The proposed model favors groups that are dominant in each language's culture, indicating bias amplification, after multilingual finetuning.
Global Voices, Local Biases: Socio-Cultural Prejudices across Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on human biases are heavily skewed towards Western and European languages . despite growing interest in language models, there are several shortcomings in the literature .
Approach: They scale the Word Embedding Association Test to 24 languages and add culturally relevant information for each language.
Outcome: The proposed language models can reflect and often amplify the effects of bias across linguistic, cultural, and societal borders.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations