Papers by Derek Powell
TAXI: Evaluating Categorical Knowledge Editing for Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Knowledge editing aims to inject new facts into language models to improve factuality, but current benchmarks fail to evaluate consistency, which is critical to ensure efficient, accurate, and generalizable edits. |
| Approach: | They manually create a new benchmark dataset specifically created to evaluate consistency in categorical knowledge edits. |
| Outcome: | The results show that the editors achieve marginal, yet non-random consistency, and their consistency far underperforms human baselines. |
Evaluating Large Language Models for Belief Inference: Mapping Belief Networks at Scale (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Beliefs are interconnected, influencing how people process and update what they think. |
| Approach: | They propose to use a finetuned GPT-4o model to infer belief structures from large-scale social media data. |
| Outcome: | The proposed model can recover belief structures from large social media data, allowing for a level of scalability and efficiency that is impossible using traditional survey methods. |