Challenge: Existing methods to measure stereotypes in large language models rely on manual templates or natural sentences that contain stereotypes.
Approach: They propose a prompt-based method to measure stereotypes in large language models . they use natural language descriptions of the target demographic group alongside unmarked defaults .
Outcome: The proposed method detects that portrayals contain higher rates of racial stereotypes than human-written portrayals.

Similar Papers

Simulating Identity, Propagating Bias: Abstraction and Stereotypes in LLM-Generated Text (2025.findings-emnlp)

Copied to clipboard

Challenge: Persona-prompting is a growing strategy to personalize outputs, but its impact on how LLMs represent social groups remains underexplored.
Approach: They investigate whether persona-prompting leads to different levels of linguistic abstraction . they compare 11 persona driven responses to those of a generic AI assistant .
Outcome: The proposed method can be used to personalize outputs, but its impact on how LLMs represent social groups remains underexplored.
The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: persona prompting is increasingly used in large language models to simulate views of various sociodemographic groups.
Approach: They use open-source LLMs to study how persona prompts influence LLM simulations . they use role adoption formats and demographic priming strategies to study marginalized groups .
Outcome: The results show that the choice of demographic priming and role adoption strategy significantly impacts their portrayal.
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans .
Approach: They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs.
Outcome: The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs.
Reading Between the Prompts: How Stereotypes Shape LLM’s Implicit Personalization (2025.emnlp-main)

Copied to clipboard

Challenge: Prior work has shown that such inferences can lead to lower quality responses for users assumed to be from minority groups.
Approach: They analyze LLMs' latent user representations through both model internals and generated answers to targeted user questions.
Outcome: The proposed models infer demographic attributes based on stereotypical signals, which persists even when the user explicitly identifies with a different demographic group.
Which Demographics do LLMs Default to During Annotation? (2025.acl-long)

Copied to clipboard

Challenge: Demographics and cultural background of annotators influence the labels they assign in text annotation.
Approach: They examine the attributes of human annotators LLMs inherently mimic and compare them to demographic-conditioned prompts and placebo-conditioned ones.
Outcome: The proposed model incorporates demographics and cultural background into the output of the large language models (LLMs) to evaluate which attributes of human annotators LLMs inherently mimic.
Persona Prompting as a Lens on LLM Social Reasoning (2026.eacl-long)

Copied to clipboard

Challenge: Persona prompting (PP) is increasingly used to steer large language models towards user-specific generation, but its effect on rationales remains underexplored.
Approach: They examine how LLM-generated rationales vary when conditioned on different demographic personas . they use word-level rationale annotations to measure agreement with human annotations based on PP .
Outcome: The proposed model improves classification on the most subjective task, but fails to align with real-world demographic counterparts.
Quantifying the Persona Effect in LLM Simulations (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown remarkable promise in simulating human language and behavior.
Approach: They investigate how integrating persona variables—demographic, social, and behavioral factors—impacts LLMs’ ability to simulate diverse perspectives.
Outcome: The proposed model improves on a zero-shot model with persona prompting.
Auditing LLM Responses to Harmful Stereotypes Targeting Mental Health Groups (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) can exhibit imbalanced biases against vulnerable groups, but how they rationalize stereotypes and rights restrictions targeting mental health entities remains underexplored.
Approach: They audit a suite of open-weight LLMs on stereotype-justification prompts tied to mental health identities.
Outcome: The proposed models endorse harmful stereotypes when explicitly asked to justify them, with endorsement varying across model families, versions, and mental health conditions.
LLM Tropes: Revealing Fine-Grained Values and Opinions in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to evaluate latent values and opinions in large language models suffer from three notable shortcomings.
Approach: They propose to analyze 156k LLM responses to 62 propositions of the Political Compass Test (PCT) generated by 6 LLMs using 420 prompt variations.
Outcome: The proposed analysis of 156k LLM responses to the Political Compass Test (PCT) generated by 6 LLMs shows that tropes are recurrent and consistent across prompts.
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing research on stereotypes in large language models is limited and focuses on African Ameri- F.
Approach: They propose to use global bias to probe a set of large language models via perplexity to determine how certain stereotypes are represented in the model's internal representations.
Outcome: The proposed model amplifys harmful stereotypes and shows that the demographic groups associated with stereotypes remain consistent across model likelihoods and outputs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations