Papers by Peiqi Sui
KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning (2025.acl-long)
Copied to clipboard
Peiqi Sui, Juan Diego Rodriguez, Philippe Laban, J. Dean Murphy, Joseph P. Dexter, Richard Jean So, Samuel Baker, Pramit Chaudhuri
| Challenge: | a study of close reading skills in large language models (LLMs) shows that LLMs still lag behind human evaluators on 10 of 11 tasks. |
| Approach: | They propose a benchmark to evaluate close reading skills in large language models . they propose three tasks to approximate different elements of the close reading process . |
| Outcome: | The proposed benchmarks show that state-of-the-art LLMs possess some college-level close reading competency, but performance still trails human evaluators on 10 out of 11 tasks. |
Confabulation: The Surprising Value of Large Language Model Hallucinations (2024.acl-long)
Copied to clipboard
| Challenge: | 'confabulations' are inherently problematic and AI research should eliminate this flaw, but confabulation is not a problem. |
| Approach: | They argue that measurable semantic characteristics of large language model (LLM) hallucinations mirror a human propensity to utilize increased narrativity as a cognitive resource for sense-making and communication. |
| Outcome: | The proposed study shows that measurable semantic characteristics of LLM confabulations mirror human propensity to utilize increased narrativity as a cognitive resource for sense-making and communication. |