Investigating Gender Stereotypes in Large Language Models via Social Determinants of Health (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing benchmarks evaluate biases related to individual social determinants of health (SDoH) but they overlook interactions between these factors and lack context-specific assessments. |
| Approach: | They investigated the relationship between gender and other SDoH in french patient records to determine whether LLMs rely on embedded stereotypes to make gendered decisions. |
| Outcome: | The proposed models can probe stereotypes and make gendered decisions based on the data. |
Similar Papers
Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes. |
| Approach: | They use a decision-making lens to examine gender equity within large language models . they explore relationships through typical and gender-neutral names . |
| Outcome: | The proposed model generation and classification models exhibit stereotypical gender biases . the proposed model generates gender-neutral names, with and without safety enhancements, and egalitarian versus traditional scenarios across topics. |
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing research on stereotypes in large language models is limited and focuses on African Ameri- F. |
| Approach: | They propose to use global bias to probe a set of large language models via perplexity to determine how certain stereotypes are represented in the model's internal representations. |
| Outcome: | The proposed model amplifys harmful stereotypes and shows that the demographic groups associated with stereotypes remain consistent across model likelihoods and outputs. |
How Can We Diagnose and Treat Bias in Large Language Models for Clinical Decision-Making? (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent studies have shown that LLMs exhibit social biases inherited from training data. |
| Approach: | They propose a framework for evaluation and mitigation of bias in Large Language Models applied to complex clinical cases using a dataset based on the JAMA Clinical Challenge. |
| Outcome: | The proposed framework employs multiple choice questions and explanations to evaluate gender and ethnicity biases in LLMs. |
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans . |
| Approach: | They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs. |
| Outcome: | The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs. |
Ask LLMs Directly, “What shapes your bias?”: Measuring Social Bias in Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to evaluate social bias in large language models have limitations . et al., 1995: stereotypes shape social perceptions without objective basis . |
| Approach: | They propose a method to intuitively quantify social perceptions and suggest metrics to evaluate biases within LLMs. |
| Outcome: | The proposed metrics capture the multi-dimensional aspects of social bias, the paper shows . they show that the proposed metrics can be used to evaluate bias in large language models . |
Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution (2024.acl-long)
Copied to clipboard
| Challenge: | Large language models reflect societal norms and biases, especially about gender. |
| Approach: | They propose to use large language models to examine gendered emotion attribution in five state-of-the-art LLMs to investigate whether emotions are genderes and whether they are influenced by societal stereotypes. |
| Outcome: | The proposed models exhibit gendered emotions, influenced by gender stereotypes, and the results are consistent with established research in psychology and gender studies. |
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models (2025.acl-short)
Copied to clipboard
| Challenge: | Large language models (LLMs) rely on superficial cues leading to spurious predictions . recent work has highlighted how LLMs exploit spurious patterns rather than learning causal, generalizable features. |
| Approach: | They use a social history annotation corpus dataset to examine drug status extraction . they evaluate prompt engineering and chain-of-thought reasoning to reduce false positives . |
| Outcome: | The proposed model can predict drug use when alcohol or smoking is not present, while uncovering gender disparities in model performance. |
“You Gotta be a Doctor, Lin” : An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated racial and gender biases in various applications. |
| Approach: | They use Large Language Models to simulate hiring decisions and salary recommendations for candidates with 320 first names that strongly signal their race and gender, across over 750,000 prompts. |
| Outcome: | The proposed models favor candidates with White female-sounding names over other demographic groups across 40 occupations. |
Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective (2025.coling-main)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) are not effective in solving real-world healthcare tasks, but they are able to provide demographic information and provide biased health predictions. |
| Approach: | They evaluate state-of-the-art LLMs with three prevalent learning frameworks across six diverse healthcare tasks and find significant challenges in applying LLM to real-world healthcare tasks. |
| Outcome: | The proposed models perform poorly in real-world healthcare tasks and are inconsistent with existing learning frameworks. |
LLMs Reproduce Stereotypes of Sexual and Gender Minorities (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a large body of research has found substantial gender bias in NLP systems . authors show that LLMs generate stereotyped representations of sexual and gender minorities in this setting . |
| Approach: | They propose to use a stereotype content model to study gender bias in large language models . they show that LLMs generate stereotyped representations of sexual and gender minorities . |
| Outcome: | The proposed model generates negative stereotypes of sexual and gender minorities in English-language surveys . |