Papers with men
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Representative bias is a tendency of Large Language Models to generate outputs that mirror the experiences of certain identity groups, and affinity bias is an evaluative preference for specific narratives. |
| Approach: | They propose two new metrics to measure representative bias and affinity bias within large language models and present a new set of tasks designed with customized rubrics to detect these biases. |
| Outcome: | The proposed model identifies representative biases in prominent LLMs, with a preference for identities associated with being white, straight, and men. |
Gendered Mental Health Stigma in Masked Language Models (2022.emnlp-main)
Copied to clipboard
Inna Lin, Lucille Njoo, Anjalie Field, Ashish Sharma, Katharina Reinecke, Tim Althoff, Yulia Tsvetkov
| Challenge: | Mental health stigma prevents many individuals from receiving appropriate care, and social psychology studies have shown that mental health tends to be overlooked in men. |
| Approach: | They propose to use clinical psychology literature to curate prompts, then evaluate models’ propensity to generate gendered words. |
| Outcome: | The proposed framework captures stigma about gender in mental health and is more likely to predict female subjects than male in sentences about mental health conditions (32% vs. 19%), and this disparity is exacerbated for sentences that indicate treatment-seeking behavior. |
Unsupervised Discovery of Gendered Language through Latent-Variable Modeling (P19-1)
Copied to clipboard
| Challenge: | a recent study has focused on the ways in which language is gendered . positive adjectives used to describe women are more often related to their bodies . |
| Approach: | They propose a model that models adjective choice and its sentiment given the natural gender of a head noun. |
| Outcome: | The proposed model shows that positive adjectives used to describe women are more often related to their bodies than positive adjective words used to explain men. |
SOLO: A Corpus of Tweets for Examining the State of Being Alone (2020.lrec-1)
Copied to clipboard
| Challenge: | Psychologists distinguish between the concept of solitude, a positive state of voluntary aloneness, and the concept 'loneliness', characterized as dissatisfaction with the quality of one’s social interactions. |
| Approach: | They present a corpus of over 4 million tweets with query terms solitude, lonely, and loneliness. |
| Outcome: | The proposed analysis analyzes over 4 million tweets with the terms solitude, lonely, and loneliness. |
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Algorithmic fairness has traditionally adopted the mathematically convenient perspective of racial color-blindness. |
| Approach: | They propose a benchmark suite of eight different scenarios to assess group difference awareness. |
| Outcome: | The proposed model demonstrates that group difference awareness is a distinct dimension to fairness where existing bias mitigation strategies may backfire. |
Multimodal Conversation Structure Understanding (2026.eacl-long)
Copied to clipboard
| Challenge: | a new set of tasks is being developed to parse the structure of conversation . female characters are 1.2 times more likely to be cast as an addressee or side-participant . |
| Approach: | They propose a set of tasks and release an annotated dataset for multimodal conversation structure understanding. |
| Outcome: | The proposed model outperforms the baseline model, but performance drops when character identities are anonymized. |
Self Promotion in US Congressional Tweets (2021.naacl-main)
Copied to clipboard
| Challenge: | Prior studies have found that women self-promote less than men due to gender stereotypes. |
| Approach: | They built a BERT-based NLP model to predict whether a Congressional tweet shows self-promotion and then used it to examine whether he gender gap exists among Congressional Tweets. |
| Outcome: | The model predicts whether a Congressional tweet shows self-promotion and then tests it against 2 million tweets from 2017 to 2021. |
Automatically Inferring Gender Associations from Language (D19-1)
Copied to clipboard
| Challenge: | In this paper, we demonstrate that there are large-scale differences in the ways that people talk about women and men and that these differences vary across domains. |
| Approach: | They propose to integrate two datasets and a novel approach to automatically infer gender associations from language and find coherent word clusters and label clusters for the semantic concepts they represent. |
| Outcome: | The proposed methods outperform strong baselines in large-scale studies of how people talk about women and men in two different settings. |
Reducing Gender Bias in Neural Machine Translation as a Domain Adaptation Problem (2020.acl-main)
Copied to clipboard
| Challenge: | Training data for NLP tasks often exhibits gender bias in that fewer sentences refer to women than to men. |
| Approach: | They propose a lattice-rescoring scheme which allows a trade-off between general translation quality and bias reduction during adaptation and inference time. |
| Outcome: | The proposed approach outperforms all systems evaluated on WinoMT with no degradation of general test set BLEU. |
LLMs Reproduce Stereotypes of Sexual and Gender Minorities (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a large body of research has found substantial gender bias in NLP systems . authors show that LLMs generate stereotyped representations of sexual and gender minorities in this setting . |
| Approach: | They propose to use a stereotype content model to study gender bias in large language models . they show that LLMs generate stereotyped representations of sexual and gender minorities . |
| Outcome: | The proposed model generates negative stereotypes of sexual and gender minorities in English-language surveys . |
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi (2025.emnlp-main)
Copied to clipboard
| Challenge: | Honorifics encode nuanced socio-pragmatic cues such as power, age, gender, fame, and cultural distance. |
| Approach: | They propose to study third-person honorific usage across 10,000 Hindi and Bengali Wikipedia articles . honorifics are more prevalent in Bengali than in Hindi, while non-honorifics dominate . |
| Outcome: | The authors show that large language models internalize similar socio-pragmatic norms . their analysis shows that honorifics are more prevalent in Bengali than in Hindi . |
If Eleanor Rigby Had Met ChatGPT: A Study on Loneliness in a Post-LLM World (2025.acl-long)
Copied to clipboard
| Challenge: | Loneliness is a global health concern and is prevalent worldwide . |
| Approach: | They analysed user interactions with ChatGPT outside of its marketed use as a task-oriented assistant and found that LLMs are more prevalent and riskier than LLM-based services . |
| Outcome: | The proposed models modify the LLMs to respond to loneliness and provide better engagement in conversations. |
EuroGEST: Investigating gender stereotypes in multilingual language models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models encode social biases, but most benchmarks for gender bias remain English-centric. |
| Approach: | They propose a dataset to measure gender-stereotypical reasoning in large language models across English and 29 European languages. |
| Outcome: | The proposed method is highly accurate across languages and strong in translations and gender labels. |