Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race (2025.acl-long)
Copied to clipboard
| Challenge: | et al., 2012) show value-aligned language models exhibit stereotypes in word association tasks . ignoring racial nuances can perpetuate subtle biases in LMs . |
| Approach: | They propose a bias mitigation strategy that incentivizes representation of racial concepts in early model layers. |
| Outcome: | The proposed approach incentivizes representation of racial concepts in early model layers . it reduces implicit bias by reducing the number of ambiguous inputs, the authors show . |
Similar Papers
One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing studies suggest large language models can capture certain behavioral patterns, but there are ongoing debates as to whether they are valid replacements for human subjects. |
| Approach: | They propose to use large language models as replacements for humans in behavioral research by relating the internal variability of simulated individuals to the population-level variability. |
| Outcome: | The proposed model can capture human-like conceptual diversity, but it is unclear whether post-training alignment affects models’ internal diversity. |
Language Models Resist Alignment: Evidence From Data Compression (2025.acl-long)
Copied to clipboard
Jiaming Ji, Kaile Wang, Tianyi Alex Qiu, Boyuan Chen, Jiayi Zhou, Changye Li, Hantao Lou, Josef Dai, Yunhuai Liu, Yaodong Yang
| Challenge: | Large language models (LLMs) may exhibit undesirable behaviors due to the inevitable biases and harmful content present in training. |
| Approach: | They propose to investigate the elasticity of large language models by examining their performance. |
| Outcome: | The proposed model performance declines rapidly before reverting to the pre-training distribution, the authors show . the proposed model weight and code are available at pku-lm-res ist-alignment.github.io. |
Aligning Language Models to Explicitly Handle Ambiguity (2024.emnlp-main)
Copied to clipboard
Hyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, Taeuk Kim
| Challenge: | Large language models (LLMs) are not specifically trained to deal with ambiguous utterances . ambiguity can lead to varying interpretations of the same input based on different assumptions or background knowledge . |
| Approach: | They propose a pipeline that aligns large language models to manage ambiguous queries . they propose to use their own assessment of perceived ambiguity to detect and manage queries a . |
| Outcome: | Experimental results show that APA empowers LLMs to detect and manage ambiguous queries while retaining the ability to answer clear questions. |
Addressing Healthcare-related Racial and LGBTQ+ Biases in Pretrained Language Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) propagate social stigmas and stereotypes, a critical concern given their widespread use. |
| Approach: | They adapt two intrinsic bias benchmarks to quantify racial and LGBTQ+ biases in prevalent PLMs and empirically evaluate the effectiveness of various debiasing methods in mitigating these biase. |
| Outcome: | The proposed methods reduce biases without compromising performance in downstream tasks. |
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection (2026.acl-long)
Copied to clipboard
| Challenge: | censored models outperform uncensoreed counterparts in accuracy and robustness, achieving 69.0% accuracy versus 64.1% strict accuracy. |
| Approach: | They examine how large language models with minimal safety alignment compare with more heavily aligned counterparts when deployed using political personas. |
| Outcome: | The proposed model outperforms uncensored models in accuracy and robustness, while uncensors are more malleable to ideological framing. |
The Authors Matter: Understanding and Mitigating Implicit Bias in Deep Text Classification (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on text classification have focused on the bias towards the individuals mentioned in the text content. |
| Approach: | They propose a framework to mitigate implicit bias in text classification models based on demographic attributes of authors . they propose to use this framework to train deep text classifiers to make predictions on the right features . |
| Outcome: | The proposed framework outperforms existing models significantly in fairness and performance. |
BiasDora: Exploring Hidden Biased Associations in Vision-Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on social biases focus on a limited set of documented associations, such as gender-profession or race-crime. |
| Approach: | They propose to examine hidden, implicit bias associations across 9 bias dimensions by probing VLMs to uncover hidden, unexamined associations. |
| Outcome: | The proposed methods reveal that biases vary in negativity, toxicity, and extremity. |
Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback (2025.findings-emnlp)
Copied to clipboard
| Challenge: | popular training paradigms for language models often assume there is one optimal answer for every query. |
| Approach: | They propose to enhance pluralistic alignment of language models using pluralistic decoding and model steering methods. |
| Outcome: | The proposed methods improve pluralistic alignment of language models in a low-resource setting . the proposed methods decrease false positives in several high-stakes tasks . |
Veracity Bias and Beyond: Uncovering LLMs’ Hidden Beliefs in Problem-Solving Reasoning (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been aligned to avoid harmful biases and stereotypes, but recent studies have revealed the superficial nature of this alignment. |
| Approach: | They propose to use large language models to avoid harmful biases and stereotypes by assigning personas to LLMs to observe decision discrepancies in social scenarios or asking them to associate specific attributes with social targets. |
| Outcome: | The proposed models attribute fewer correct solutions and more incorrect ones to African-American groups in math and coding, while Asian authorships are least preferred in writing evaluation. |
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models can lead to undesired consequences when misaligned with human values . previous studies have shown misalignment of LLMs with human value using expert-designed or agent-based emulated bias scenarios . |
| Approach: | They investigate whether large language models (LLMs) are misaligned with human values . they find no significant differences in understanding of HVSB between LLMs . |
| Outcome: | The results show that large language models do not have lower misalignment rates and attack success rates . the study also shows that smaller language models have the ability to explain HVSB . |