Papers by Saloni Dash
Anecdoctoring: Automated Red-Teaming Across Language and Place (2025.emnlp-main)
Copied to clipboard
| Challenge: | Disinformation is among the top risks of generative AI misuse . red-teaming datasets are typically US- and English-centric . |
| Approach: | They propose a red-teaming approach that generates adversarial prompts across languages and cultures by clustering misinformation claims into broader narratives and enhancing an attacker LLM. |
| Outcome: | The proposed approach produces higher attack success rates and interpretability benefits relative to few-shot prompting. |
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior studies have reported that large language models (LLMs) are also susceptible to human-like cognitive biases, but the extent to which LLMs selectively reason toward identity-congruent conclusions remains unexplored. |
| Approach: | They investigate whether assigning 8 personas across 4 political and socio-demographic attributes induces motivated reasoning in LLMs. |
| Outcome: | The proposed model is assigned 8 personas across 4 political and socio-demographic attributes and shows that they have 9% reduced veracity discernment compared to models without persona. |