Papers by Maziar Raissi
Aligning to What? Limits to RLHF Based Alignment (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies on RLHF and covert and overt biases in large language models are unclear . et al. analyzed off-the-shelf language models to evaluate their overt and cover racial biase . |
| Approach: | They evaluate the relationship between reinforcement learning from human feedback and biases in large language models. |
| Outcome: | The proposed approach can be used to mitigat covert biases, the authors show . they found that the RLHF approach calcifies model biase . |