Challenge: Large Language Models (LLMs) can exhibit imbalanced biases against vulnerable groups, but how they rationalize stereotypes and rights restrictions targeting mental health entities remains underexplored.
Approach: They audit a suite of open-weight LLMs on stereotype-justification prompts tied to mental health identities.
Outcome: The proposed models endorse harmful stereotypes when explicitly asked to justify them, with endorsement varying across model families, versions, and mental health conditions.

Similar Papers

Rethinking Research on Stereotypes: An Analysis through Social Psychological and Computational Perspectives (2026.findings-acl)

Copied to clipboard

Challenge: Existing research on stereotypical biases ignores literature on them and results in resource wastage.
Approach: They argue that stereotypes are social constructs shaping human perception and behavior that can produce harmful outcomes under specific conditions.
Outcome: The proposed models can inherit and amplify stereotypes under certain conditions.
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings.
Approach: They propose safety guidelines for the potential deployment of large language models for mental health response.
Outcome: The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory.
LLM Tropes: Revealing Fine-Grained Values and Opinions in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to evaluate latent values and opinions in large language models suffer from three notable shortcomings.
Approach: They propose to analyze 156k LLM responses to 62 propositions of the Political Compass Test (PCT) generated by 6 LLMs using 420 prompt variations.
Outcome: The proposed analysis of 156k LLM responses to the Political Compass Test (PCT) generated by 6 LLMs shows that tropes are recurrent and consistent across prompts.
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans .
Approach: They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs.
Outcome: The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs.
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications (2024.naacl-long)

Copied to clipboard

Challenge: Recent studies suggest using large language models to make tabular classifications . however, LLMs have been shown to exhibit harmful social biases based on stereotypes and inequalities present in society.
Approach: They propose to use large language models to make tabular classifications . they show that LLMs inherit biases from their training data .
Outcome: The proposed models exhibit harmful biases that reflect stereotypes and inequalities in society.
Investigating Gender Stereotypes in Large Language Models via Social Determinants of Health (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks evaluate biases related to individual social determinants of health (SDoH) but they overlook interactions between these factors and lack context-specific assessments.
Approach: They investigated the relationship between gender and other SDoH in french patient records to determine whether LLMs rely on embedded stereotypes to make gendered decisions.
Outcome: The proposed models can probe stereotypes and make gendered decisions based on the data.
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have also introduced additional safety risks and raised concerns regarding their detrimental impact on already marginalized populations.
Approach: They propose to use LLMs to evaluate their safety responses on already mitigated biases by evaluating models on already encoded assumptions.
Outcome: The proposed model can encode harmful assumptions, but it can also be harmful for certain demographic groups.
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs (2025.findings-acl)

Copied to clipboard

Challenge: 211 studies on the demographic representativeness of large language models have conflicting results . 29% of the studies report positive conclusions on the representativeness, 30% do not evaluate LLMs across multiple demographic categories or within demographic subcategories.
Approach: 211 papers review the representativeness of large language models . authors recommend more precise evaluation methods and comprehensive documentation of demographic attributes .
Outcome: 211 studies on the representativeness of large language models are reviewed . 29% of the studies report positive conclusions, but 30% fail to specify subcategories . authors recommend more precise evaluation methods and documentation of demographic attributes .
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies have shown that relying on LLMs as information providers may hurt student learning.
Approach: They introduce and apply two bias score metrics to evaluate LLMs for bias in the personalized educational setting, specifically on the models’ roles as “teachers.”
Outcome: The proposed models harm student learning by perpetuating harmful stereotypes and reversing them.
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing research on stereotypes in large language models is limited and focuses on African Ameri- F.
Approach: They propose to use global bias to probe a set of large language models via perplexity to determine how certain stereotypes are represented in the model's internal representations.
Outcome: The proposed model amplifys harmful stereotypes and shows that the demographic groups associated with stereotypes remain consistent across model likelihoods and outputs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations