Papers by Xuechunzi Bai
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race (2025.acl-long)
Copied to clipboard
| Challenge: | et al., 2012) show value-aligned language models exhibit stereotypes in word association tasks . ignoring racial nuances can perpetuate subtle biases in LMs . |
| Approach: | They propose a bias mitigation strategy that incentivizes representation of racial concepts in early model layers. |
| Outcome: | The proposed approach incentivizes representation of racial concepts in early model layers . it reduces implicit bias by reducing the number of ambiguous inputs, the authors show . |