Papers by Cassandra Overney
On the Relationship between Truth and Political Bias in Language Models (2024.emnlp-main)
Copied to clipboard
Suyash Fulay, William Brannon, Shrestha Mohanty, Cassandra Overney, Elinor Poole-Dayan, Deb Roy, Jad Kabbara
| Challenge: | Language model alignment research often attempts to ensure that models are helpful and harmless, but can obscure how improving one aspect might impact the other. |
| Approach: | They analyze the relationship between truthfulness and political bias in language models. |
| Outcome: | The results show that optimizing models for truthfulness results in a left-leaning political bias. |