French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English (2022.acl-long)
Copied to clipboard
| Challenge: | We introduce 1,679 sentence pairs in French that cover stereotypes in ten types of bias like gender and age. |
| Approach: | They build on the US-centered CrowS-pairs dataset to create a multilingual stereotypes dataset that allows for comparability across languages and cultures. |
| Outcome: | The proposed dataset allows for comparability across languages while characterizing biases that are specific to each country and language. |
Similar Papers
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models (2020.emnlp-main)
Copied to clipboard
| Challenge: | Pretrained language models use cultural biases implicitly, causing harm . identifying and quantifying learnt biase enables us to measure progress . |
| Approach: | They propose a benchmark to measure social bias in pretrained language models . they use 1508 examples that cover stereotypes dealing with nine types of bias . |
| Outcome: | The proposed benchmark focuses on stereotypes about historically disadvantaged groups and contrasts them with advantaged groups. |
Your Stereotypical Mileage May Vary: Practical Challenges of Evaluating Biases in Multiple Languages and Cultural Contexts (2024.lrec-main)
Copied to clipboard
Karen Fort, Laura Alonso Alemany, Luciana Benotti, Julien Bezançon, Claudia Borg, Marthese Borg, Yongjian Chen, Fanny Ducel, Yoann Dupont, Guido Ivetta, Zhijian Li, Margot Mieskes, Marco Naguib, Yuyan Qian, Matteo Radaelli, Wolfgang S. Schmeisser-Nieto, Emma Raimundo Schulz, Thiziri Saci, Sarah Saidi, Javier Torroba Marchante, Shilin Xie, Sergio E. Zanotto, Aurélie Névéol
| Challenge: | Recent studies have identified a gap in the availability of tools and resources to study bias in languages other than English and social contexts outside the north of America. |
| Approach: | They use stereotypes to build a corpus of sentence pairs that cover biases in seven cultural contexts. |
| Outcome: | The proposed resource covers a wide range of languages and cultural settings . it favors sentences that express stereotypes in most bias categories . |
SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models (2025.naacl-long)
Copied to clipboard
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, null Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
| Challenge: | Large Language Models reproduce and exacerbate social biases present in training data, and resources to quantify this issue are limited. |
| Approach: | They propose a multilingual parallel dataset to examine culturally-specific stereotypes that may be learned by LLMs. |
| Outcome: | The proposed dataset includes stereotypes from 20 regions around the world and 16 languages, spanning multiple identity categories subject to discrimination worldwide. |
An Information-Theoretic Approach and Dataset for Probing Gender Stereotypes in Multilingual Masked Language Models (2022.findings-naacl)
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) have been shown to encapsulate social biases, including those relating to gender and race. |
| Approach: | They propose a new bias measure based on Jensen–Shannon divergence that retains more information from the model output probabilities than other previously proposed bias measures. |
| Outcome: | The proposed measure outperforms CrowS-Pairs and other similar measures for non-English datasets. |
Social Bias in Multilingual Language Models: A Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Pretrained multilingual models exhibit the same social bias as models processing English texts. |
| Approach: | They examine the literature on bias evaluation and mitigation approaches in multilingual and non-English contexts and identify gaps in the field. |
| Outcome: | The proposed models perform well on multilingual language understanding benchmarks and are consistent with the current literature. |
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing research on stereotypes in large language models is limited and focuses on African Ameri- F. |
| Approach: | They propose to use global bias to probe a set of large language models via perplexity to determine how certain stereotypes are represented in the model's internal representations. |
| Outcome: | The proposed model amplifys harmful stereotypes and shows that the demographic groups associated with stereotypes remain consistent across model likelihoods and outputs. |
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks for measuring gender stereotypical bias in language models are inconsistencies . lack of explicit standards in data gathering can have detrimental effects on results . |
| Approach: | They propose that currently available benchmarks capture only partial facets of gender stereotypes . they apply a framework from social psychology to balance data across components of gender stereotypes based on stereotypical benchmarks. |
| Outcome: | The proposed framework improves correlation between different benchmarks by using simple balancing techniques. |
In-Depth Look at Word Filling Societal Bias Measures (2023.eacl-main)
Copied to clipboard
| Challenge: | Language models (LMs) are ubiquitous in current NLP and have brought undeniable performance improvements for many tasks. |
| Approach: | They propose to use word filling prompts to evaluate language models' behavior to find out if they are valid. |
| Outcome: | The proposed measures produce unexpected and illogical results when appropriate control group samples are constructed. |
Comparing Biases and the Impact of Multilingual Training across Multiple Languages (2023.emnlp-main)
Copied to clipboard
Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, Dan Roth
| Challenge: | Currently, studies on bias and fairness in natural language processing focus on a single language and/or across few attributes (e.g. gender, race). However, biases can manifest differently across languages for individual attributes. |
| Approach: | They adapt existing sentiment bias templates in English to Italian, Chinese, Hebrew, and Spanish for race, religion, nationality, and gender. |
| Outcome: | The proposed model favors groups that are dominant in each language's culture, indicating bias amplification, after multilingual finetuning. |
Global Voices, Local Biases: Socio-Cultural Prejudices across Languages (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on human biases are heavily skewed towards Western and European languages . despite growing interest in language models, there are several shortcomings in the literature . |
| Approach: | They scale the Word Embedding Association Test to 24 languages and add culturally relevant information for each language. |
| Outcome: | The proposed language models can reflect and often amplify the effects of bias across linguistic, cultural, and societal borders. |