Papers by Sil Hamilton
Fiction Flows: A Replication and Reinterpretation of Narrative Sequentiality (2026.acl-long)
Copied to clipboard
| Challenge: | a new study shows that imagined narratives exhibit higher "flow" than recalled narratives, but this advantage is not reducible to standard coherence measures. |
| Approach: | They propose a language-model-based measure of sentence-level predictability to measure narrative flow . they find that imagined stories flow better than recalled ones . |
| Outcome: | The proposed measure of sentence-level predictability is based on language models . it shows that fiction exhibits a robust sequentiality advantage over reality-bound genres . |
NarraBench: A Comprehensive Framework for Narrative Benchmarking (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for narrative understanding are poorly aligned with existing metrics. |
| Approach: | They propose to use NarraBench to assess aspects of narrative understanding that are either overlooked in current work or are poorly aligned with existing metrics. |
| Outcome: | The proposed taxonomy and survey are useful to NLP researchers . they find that only 27% of tasks are well captured by existing benchmarks . |
SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models (2025.naacl-long)
Copied to clipboard
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, null Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
| Challenge: | Large Language Models reproduce and exacerbate social biases present in training data, and resources to quantify this issue are limited. |
| Approach: | They propose a multilingual parallel dataset to examine culturally-specific stereotypes that may be learned by LLMs. |
| Outcome: | The proposed dataset includes stereotypes from 20 regions around the world and 16 languages, spanning multiple identity categories subject to discrimination worldwide. |