Papers by Peiqi Sui

2 papers
KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning (2025.acl-long)

Copied to clipboard

Challenge: a study of close reading skills in large language models (LLMs) shows that LLMs still lag behind human evaluators on 10 of 11 tasks.
Approach: They propose a benchmark to evaluate close reading skills in large language models . they propose three tasks to approximate different elements of the close reading process .
Outcome: The proposed benchmarks show that state-of-the-art LLMs possess some college-level close reading competency, but performance still trails human evaluators on 10 out of 11 tasks.
Confabulation: The Surprising Value of Large Language Model Hallucinations (2024.acl-long)

Copied to clipboard

Challenge: 'confabulations' are inherently problematic and AI research should eliminate this flaw, but confabulation is not a problem.
Approach: They argue that measurable semantic characteristics of large language model (LLM) hallucinations mirror a human propensity to utilize increased narrativity as a cognitive resource for sense-making and communication.
Outcome: The proposed study shows that measurable semantic characteristics of LLM confabulations mirror human propensity to utilize increased narrativity as a cognitive resource for sense-making and communication.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations