Challenge: We study whether large language models exhibit race- and gender-based name discrimination in hiring decisions .
Approach: They propose templatic prompts to LLMs to write an email to a named job applicant informing them of a hiring decision.
Outcome: The proposed model generates an acceptance or rejection email based on the applicant's first name .

Similar Papers

“You Gotta be a Doctor, Lin” : An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated racial and gender biases in various applications.
Approach: They use Large Language Models to simulate hiring decisions and salary recommendations for candidates with 320 first names that strongly signal their race and gender, across over 750,000 prompts.
Outcome: The proposed models favor candidates with White female-sounding names over other demographic groups across 40 occupations.
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have potential to automate hiring but inherent biases may lead to unfair hiring practices.
Approach: They evaluate how factors such as gender, race, and educational background influence model decisions.
Outcome: The proposed model reduces biases related to gender and race, but implicit biase concerning educational background remains significant.
Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes.
Approach: They use a decision-making lens to examine gender equity within large language models . they explore relationships through typical and gender-neutral names .
Outcome: The proposed model generation and classification models exhibit stereotypical gender biases . the proposed model generates gender-neutral names, with and without safety enhancements, and egalitarian versus traditional scenarios across topics.
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: We show that models are less likely to predict romantic relationships for same-gender character pairs than different-grace character pairs.
Approach: They perform name-replacement experiments to examine gender biases in large language models . they hypothesize that models mirror heteronormative biase and prejudice against interracial romantic relationships .
Outcome: The results suggest that models may mirror heteronormative biases and prejudice against interracial romantic relationships in human and society.
Hire Me or Not? Examining Language Model’s Behavior with Occupation Attributes (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) have been widely integrated into production pipelines due to their impressive performance across multiple tasks.
Approach: They construct a dataset using a standard occupation classification knowledge base and tested it on three families of LLMs.
Outcome: The proposed framework analyzes LLMs’ behavior with respect to gender stereotypes in the context of occupation decision making.
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: a framework for benchmarking hierarchical gender hiring bias in Large Language Models (LLMs) is developed to protect vulnerable demographic groups.
Approach: They propose a framework for benchmarking hierarchical gender hiring bias in Large Language Models for resume scoring.
Outcome: The proposed framework reveals significant issues of reverse gender hiring bias and overdebiasing in ten state-of-the-art LLMs.
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that large language models can successfully persuade humans and amplify persuasive language.
Approach: They propose a framework for evaluating how persuasive language generation is affected by recipient gender, sender intent, or output language.
Outcome: The proposed framework varies persuasive language when the recipient gender is specified or when the sender intent is specified.
On the Mutual Influence of Gender and Occupation in LLM Representations (2025.acl-long)

Copied to clipboard

Challenge: We examine LLM representations of gender for first names in various occupational contexts to study how occupations and the gender perception of first names influence each other mutually.
Approach: They examine LLM representations of gender for first names in various occupational contexts and examine how occupations and the gender perception of first names influence each other mutually.
Outcome: The representations shift with the occupational context and are influenced by stereotypically feminine or masculine occupations.
Your Mileage May Vary: How Empathy and Demographics Shape Human Preferences in LLM Responses (2025.findings-emnlp)

Copied to clipboard

Challenge: large language models (LLMs) increasingly assist subjective decision-making . prior work uses aggregate human judgments, but demographic variation and its linguistic drivers remain underexplored.
Approach: They analyze how demographic background and empathy level correlate with LLM-generated dilemma responses . they also identify markers that predict group-level differences .
Outcome: The authors show that demographic background and empathy level correlate with LLM preferences . their findings highlight the need for demographically informed LLM evaluations.
Examining Gender and Racial Bias in Large Vision–Language Models Using a Novel Dataset of Parallel Images (2024.eacl-long)

Copied to clipboard

Challenge: a new wave of large vision–language models (LVLMs) incorporate images as input in addition to text . a recent study examined potential gender and racial biases in such systems based on the perceived characteristics of the people in the input images.
Approach: They examine potential gender and racial biases in large vision–language models . they query a dataset of AI-generated images of people to see whether they differ .
Outcome: The proposed dataset shows that the images differ in gender and race according to the perceived characteristics of the person depicted.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations