Challenge: Healthcare professionals are increasingly including Language Models (LMs) in clinical practice.
Approach: They propose to use LMs to generate clinical cases in french and an automatic linguistic gender detection tool to measure gender biases.
Outcome: The proposed model over-generates cases describing male patients, creating synthetic corpora that are not consistent with documented prevalence for these disorders.

Similar Papers

Race, Gender, and Age Biases in Biomedical Masked Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models can be used to identify and eliminate healthcare disparities.
Approach: They examine social biases present in biomedical masked language models . they curate prompts based on evidence-based practice and compare generated diagnoses .
Outcome: The proposed models are less biased than BERT in gender, while the opposite is true for race and age.
How Can We Diagnose and Treat Bias in Large Language Models for Clinical Decision-Making? (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have shown that LLMs exhibit social biases inherited from training data.
Approach: They propose a framework for evaluation and mitigation of bias in Large Language Models applied to complex clinical cases using a dataset based on the JAMA Clinical Challenge.
Outcome: The proposed framework employs multiple choice questions and explanations to evaluate gender and ethnicity biases in LLMs.
Investigating Gender Stereotypes in Large Language Models via Social Determinants of Health (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks evaluate biases related to individual social determinants of health (SDoH) but they overlook interactions between these factors and lack context-specific assessments.
Approach: They investigated the relationship between gender and other SDoH in french patient records to determine whether LLMs rely on embedded stereotypes to make gendered decisions.
Outcome: The proposed models can probe stereotypes and make gendered decisions based on the data.
Evaluating Gender Bias of LLMs in Making Morality Judgements (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable capabilities in a multitude of NLP tasks, but are still not immune to limitations such as gender bias.
Approach: They propose to use a dataset to examine whether LLMs possess gender bias when asked to give moral opinions.
Outcome: The proposed models show that they are biased when asked to give moral opinions.
UnMASKed: Quantifying Gender Biases in Masked Language Models through Linguistically Informed Job Market Prompts (2024.eacl-srw)

Copied to clipboard

Challenge: Language models (LMs) often include societal biases encoded in the human-produced datasets used for their training.
Approach: They evaluated six prominent language models: BERT, RoBERTa, DistilBERT, BERT- multilingual, XLM-RoBERT and DistilberT- multilinguistic.
Outcome: The results show that the models generated by the models were stereotypically gendered and with a reduced bias in multilingual variants.
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies have shown that large language models can cause harmful, human-like biases against various demographics.
Approach: They propose a causal formulation for bias measurement in generative language models based on a list of desiderata for designing robust bias benchmarks and a bias-measuring procedure to investigate occupational gender bias.
Outcome: The proposed framework is generalizable and can be extended to include other datasets.
“Kelly is a Warm Person, Joseph is a Role Model”: Gender Biases in LLM-Generated Reference Letters (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are an effective tool to assist individuals in writing documents.
Approach: They examine gender biases in large language models (LLMs)-generated reference letters . they find that models are biased because they are hallucinated .
Outcome: The proposed model-generated reference letters are evaluated on 2 popular LLMs- ChatGPT and Alpaca.
“Fifty Shades of Bias”: Normative Ratings of Gender Bias in GPT Generated English Text (2023.emnlp-main)

Copied to clipboard

Challenge: Prior work treats gender bias as a binary classification task, but a comparative annotation framework can be used to assess the impact of biases.
Approach: They propose to generate a dataset with normative ratings of gender bias in English text with a comparative annotation framework.
Outcome: The first dataset of GPT-generated English text with normative ratings of gender bias is analyzed using Best–Worst Scaling .
Leveraging Pre-trained Language Models for Gender Debiasing (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods to reduce gender bias in natural language are costly and time-consuming.
Approach: They propose a method to generate gender variants for a given text using pre-trained language models as the resource without any task-specific labelled data.
Outcome: The proposed method can reduce gender bias in a language generation context without a task-specific labelled data.
Evaluating Gender Bias in Machine Translation (P19-1)

Copied to clipboard

Challenge: Using morphological analysis, we find that MT models exhibit gender-biased translation errors when training data encode stereotypes not relevant for the task.
Approach: They propose an automatic gender bias evaluation method for eight target languages with grammatical gender based on morphological analysis.
Outcome: The proposed method is based on two recent coreference resolution datasets composed of English sentences cast participants into non-stereotypical gender roles.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations