Auditing LLM Responses to Harmful Stereotypes Targeting Mental Health Groups (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can exhibit imbalanced biases against vulnerable groups, but how they rationalize stereotypes and rights restrictions targeting mental health entities remains underexplored. |
| Approach: | They audit a suite of open-weight LLMs on stereotype-justification prompts tied to mental health identities. |
| Outcome: | The proposed models endorse harmful stereotypes when explicitly asked to justify them, with endorsement varying across model families, versions, and mental health conditions. |
Similar Papers
Rethinking Research on Stereotypes: An Analysis through Social Psychological and Computational Perspectives (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing research on stereotypical biases ignores literature on them and results in resource wastage. |
| Approach: | They argue that stereotypes are social constructs shaping human perception and behavior that can produce harmful outcomes under specific conditions. |
| Outcome: | The proposed models can inherit and amplify stereotypes under certain conditions. |
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings. |
| Approach: | They propose safety guidelines for the potential deployment of large language models for mental health response. |
| Outcome: | The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory. |
LLM Tropes: Revealing Fine-Grained Values and Opinions in Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to evaluate latent values and opinions in large language models suffer from three notable shortcomings. |
| Approach: | They propose to analyze 156k LLM responses to 62 propositions of the Political Compass Test (PCT) generated by 6 LLMs using 420 prompt variations. |
| Outcome: | The proposed analysis of 156k LLM responses to the Political Compass Test (PCT) generated by 6 LLMs shows that tropes are recurrent and consistent across prompts. |
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans . |
| Approach: | They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs. |
| Outcome: | The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs. |
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent studies suggest using large language models to make tabular classifications . however, LLMs have been shown to exhibit harmful social biases based on stereotypes and inequalities present in society. |
| Approach: | They propose to use large language models to make tabular classifications . they show that LLMs inherit biases from their training data . |
| Outcome: | The proposed models exhibit harmful biases that reflect stereotypes and inequalities in society. |
Investigating Gender Stereotypes in Large Language Models via Social Determinants of Health (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing benchmarks evaluate biases related to individual social determinants of health (SDoH) but they overlook interactions between these factors and lack context-specific assessments. |
| Approach: | They investigated the relationship between gender and other SDoH in french patient records to determine whether LLMs rely on embedded stereotypes to make gendered decisions. |
| Outcome: | The proposed models can probe stereotypes and make gendered decisions based on the data. |
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards (2024.findings-acl)
Copied to clipboard
Khaoula Chehbouni, Megha Roshan, Emmanuel Ma, Futian Wei, Afaf Taik, Jackie Cheung, Golnoosh Farnadi
| Challenge: | Recent advances in large language models have also introduced additional safety risks and raised concerns regarding their detrimental impact on already marginalized populations. |
| Approach: | They propose to use LLMs to evaluate their safety responses on already mitigated biases by evaluating models on already encoded assumptions. |
| Outcome: | The proposed model can encode harmful assumptions, but it can also be harmful for certain demographic groups. |
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | 211 studies on the demographic representativeness of large language models have conflicting results . 29% of the studies report positive conclusions on the representativeness, 30% do not evaluate LLMs across multiple demographic categories or within demographic subcategories. |
| Approach: | 211 papers review the representativeness of large language models . authors recommend more precise evaluation methods and comprehensive documentation of demographic attributes . |
| Outcome: | 211 studies on the representativeness of large language models are reviewed . 29% of the studies report positive conclusions, but 30% fail to specify subcategories . authors recommend more precise evaluation methods and documentation of demographic attributes . |
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies have shown that relying on LLMs as information providers may hurt student learning. |
| Approach: | They introduce and apply two bias score metrics to evaluate LLMs for bias in the personalized educational setting, specifically on the models’ roles as “teachers.” |
| Outcome: | The proposed models harm student learning by perpetuating harmful stereotypes and reversing them. |
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing research on stereotypes in large language models is limited and focuses on African Ameri- F. |
| Approach: | They propose to use global bias to probe a set of large language models via perplexity to determine how certain stereotypes are represented in the model's internal representations. |
| Outcome: | The proposed model amplifys harmful stereotypes and shows that the demographic groups associated with stereotypes remain consistent across model likelihoods and outputs. |