Challenge: Existing studies have focused on measuring the degree to which pre-trained language models capture purely linguistic knowledge and reasoning abilities and world knowledge.
Approach: They use geography to demarcate different populations around the world and comparable corpora to measure how well two families of LLMs perform across these different populations.
Outcome: The results show that pre-trained models perform better for some populations than others.

Similar Papers

Measuring Geographic Performance Disparities of Offensive Language Classifiers (2022.coling-1)

Copied to clipboard

Challenge: Recent work shows that text classifiers are biased regarding different languages and dialects.
Approach: They propose to use a dataset to examine whether language, dialect, and topical content vary across geographical regions to address these gaps.
Outcome: The proposed dataset includes 14 thousand examples across 15 cities and shows that current models do not generalize across locations.
Sociolectal Analysis of Pretrained Language Models (2021.emnlp-main)

Copied to clipboard

Challenge: Using data from English cloze tests, we demonstrate wide performance gaps across demographic groups and show that pretrained language models disfavor young non-white male speakers.
Approach: They use data from English cloze tests to examine performance differences of pretrained language models across demographic groups.
Outcome: The models disfavor young non-white male speakers, but larger models reduce performance gaps between majority and minority groups.
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing research on stereotypes in large language models is limited and focuses on African Ameri- F.
Approach: They propose to use global bias to probe a set of large language models via perplexity to determine how certain stereotypes are represented in the model's internal representations.
Outcome: The proposed model amplifys harmful stereotypes and shows that the demographic groups associated with stereotypes remain consistent across model likelihoods and outputs.
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs (2026.acl-long)

Copied to clipboard

Challenge: Multilingual large language models have minimized the fluency gap between languages, but they are exposed to the risk of biases as knowledge and norms may propagate across languages.
Approach: They propose a test set with 2,156 questions in 12 languages to quantify models' biases . they show a global bias towards answers relevant to the US-locale .
Outcome: The proposed model can answer locale-ambiguous questions in 12 languages.
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have redefined Machine Translation, enabling context-aware and fluent translations across hundreds of languages and textual domains.
Approach: They propose a framework and dataset to evaluate the translation quality and fairness of open-source LLMs.
Outcome: The proposed framework and dataset evaluates translation quality and fairness of open-source LLMs.
7 Points to Tsinghua but 10 Points to ? Assessing Large Language Models in Agentic Multilingual National Bias (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models have garnered significant attention for their capabilities in multilingual natural language processing, but studies on risks associated with cross biases are limited to immediate context preferences.
Approach: They investigate multilingual bias in state-of-the-art Large Language Models by analyzing their responses to decision-making tasks across multiple languages.
Outcome: The proposed model can provide personalized advice across university applications, travel, and relocation scenarios.
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs (2025.findings-acl)

Copied to clipboard

Challenge: 211 studies on the demographic representativeness of large language models have conflicting results . 29% of the studies report positive conclusions on the representativeness, 30% do not evaluate LLMs across multiple demographic categories or within demographic subcategories.
Approach: 211 papers review the representativeness of large language models . authors recommend more precise evaluation methods and comprehensive documentation of demographic attributes .
Outcome: 211 studies on the representativeness of large language models are reviewed . 29% of the studies report positive conclusions, but 30% fail to specify subcategories . authors recommend more precise evaluation methods and documentation of demographic attributes .
StereoSet: Measuring stereotypical bias in pretrained language models (2021.acl-long)

Copied to clipboard

Challenge: Existing literature on stereotypical biases in language models is limited . current evaluations focus on measuring bias without considering language modeling ability .
Approach: They propose to measure stereotypical biases in four domains: gender, profession, race, and religion . they compare stereotypical and language modeling ability of popular models like BERT, GPT-2, RoBERTa and XLnet .
Outcome: The proposed model shows strong stereotypical biases in gender, profession, race, and religion domains.
Deciphering the Impact of Pretraining Data on Large Language Models through Machine Unlearning (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies have suggested that the composition of the pretraining corpus exerts a significant impact upon the performance of LLMs.
Approach: They analyze the impact of 48 datasets from 5 major categories of pretraining data of Large Language Models and measure their impacts on LLMs using benchmarks about nine major categories.
Outcome: The proposed analysis provides insights into the organization of data to support more efficient pretraining of Large Language Models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations