Papers by Elinor Poole-Dayan
On the Relationship between Truth and Political Bias in Language Models (2024.emnlp-main)
Copied to clipboard
Suyash Fulay, William Brannon, Shrestha Mohanty, Cassandra Overney, Elinor Poole-Dayan, Deb Roy, Jad Kabbara
| Challenge: | Language model alignment research often attempts to ensure that models are helpful and harmless, but can obscure how improving one aspect might impact the other. |
| Approach: | They analyze the relationship between truthfulness and political bias in language models. |
| Outcome: | The results show that optimizing models for truthfulness results in a left-leaning political bias. |
Computational Analysis of Conversation Dynamics through Participant Responsivity (2025.emnlp-main)
Copied to clipboard
| Challenge: | Growing literature explores toxicity and polarization in discourse, with comparatively little work on characterizing what makes dialogue prosocial and constructive. |
| Approach: | They develop and evaluate methods for quantifying responsivity through semantic similarity of speaker turns and large language models to identify the relation between two speaker turns. |
| Outcome: | The proposed method is based on semantic similarity of speaker turns and large language models to identify the relation between two speaker turns. |
An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models (2022.acl-long)
Copied to clipboard
| Challenge: | Recent work has shown pre-trained language models capture social biases from the large amounts of text they are trained on. |
| Approach: | They propose to use Counterfactual Data Augmentation, Dropout, Iterative Nullspace Projection, Self-Debias, and SentenceDebia as bias mitigation techniques to quantify their effectiveness. |
| Outcome: | The proposed techniques are Counterfactual Data Augmentation (CDA), Dropout, Iterative Nullspace Projection, Self-Debias, and SentenceDebia. |