| Challenge: | a recent study shows that long-context language models can exceed human expert performance in literary analysis . despite their speed and apparent accuracy, even the strongest models struggle with nuanced literary signals and overgeneration. |
| Approach: | They propose a task where a model is given an entire text of a book and a literary criticism with a missing quotation from that work and asked to generate the missing quote. |
| Outcome: | The proposed model outperforms open-weight models in literary evidence retrieval tasks. |
Similar Papers
Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have raised concerns about reliability and trustworthiness of the models. |
| Approach: | They analyze 134 papers and introduce a taxonomy of evidence-based text generation with LLMs. |
| Outcome: | The proposed methods highlight open challenges and outline promising directions for future work. |
Towards A “Novel” Benchmark: Evaluating Literary Fiction with Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) context windows have enabled them to process inputs over 100K tokens and generate outputs of up to 10K token. |
| Approach: | They propose a multi-level evaluation framework that incorporates ten metrics across the Macro, Meso, and Micro levels and an annotated fiction dataset. |
| Outcome: | The proposed framework incorporates ten metrics across the Macro, Meso, and Micro levels and is based on a human-human-AI dataset. |
Logic Haystacks: Probing LLMs’ Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding) (2026.eacl-short)
Copied to clipboard
| Challenge: | Recent large language models claim long context windows, but evaluations often involve simple retrieval tasks or synthetic tasks padded with irrelevant text. |
| Approach: | They use grammars to generate simplified English with logical representations to create long input text while controlling its semantics. |
| Outcome: | The proposed model performs better with realistic distractors than with standard models. |
Evaluating LLMs for Quotation Attribution in Literary Texts: A Case Study of LLaMa3 (2025.naacl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown promising results in literary tasks . however, quotation attribution remains a challenging task and methods that generalize across writing styles are lacking analysis regarding book memorization and annotation contamination. |
| Approach: | They evaluate the ability of Llama-3 to attribute utterances of direct-speech to their speaker in novels by assessing the impact of book memorization and annotation contamination. |
| Outcome: | The proposed model outperforms existing models on a corpus of 28 novels and shows that book memorization and annotation contamination do not explain the performance gain. |
DIVKNOWQA: Assessing the Reasoning Ability of LLMs via Open-Domain Question Answering over Knowledge Base and Text (2024.findings-naacl)
Copied to clipboard
| Challenge: | Retrievalaugmented LLMs have been used to ground LLM in external knowledge . a gap exists in the current landscape regarding the effectiveness of grounding LLM on heterogeneous knowledge sources. |
| Approach: | They propose a model that uses symbolic language to generate symbolic queries . they use a dataset that is generated using predefined reasoning chains and human annotation . |
| Outcome: | The proposed model outperforms previous approaches by a significant margin in QA tasks over text. |
LFED: A Literary Fiction Evaluation Dataset for Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | LFED is a literary fiction evaluation dataset for large language models that evaluate the capability of LLMs on the long fiction comprehension and reasoning. |
| Approach: | They propose a Literary Fiction Evaluation Dataset to evaluate LLMs' comprehension and reasoning on long fictions. |
| Outcome: | The proposed dataset evaluates the capability of large language models on the long fiction comprehension and reasoning. |
One Thousand and One Pairs: A “novel” challenge for long-context language models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing long-context evaluation methods measure surface-level retrieval capabilities, but do not assess performance on the more challenging task of synthesizing distant and underlying information. |
| Approach: | They propose a dataset of 1,001 minimally different pairs of true and false claims about 67 recently-published English fictional books. |
| Outcome: | The proposed model performs better on pairs that require only sentence-level retrieval vs. global reasoning . the proposed model also performs worse on speculative fiction books with extensive world-building . |
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are a promising solution to automate literature review writing tasks. |
| Approach: | They propose a framework to automatically evaluate the performance of large language models in three key tasks of literature review writing: reference generation, abstract writing, and literature review composition. |
| Outcome: | The proposed framework assesses the hallucination rates in generated references and measures the semantic coverage and factual consistency of the literature summaries and compositions against human-written counterparts. |
What Evidence Do Language Models Find Convincing? (2024.acl-long)
Copied to clipboard
| Challenge: | Current retrieval-augmented language models are tasked with subjective, contentious, and conflicting queries. |
| Approach: | They construct a dataset that pairs controversial queries with real-world evidence documents . they find current models rely heavily on relevance of a website to the query . |
| Outcome: | The proposed dataset pairs controversial queries with real-world evidence documents that contain different facts, arguments, and answers. |
Current Advances in LLM Reasoning (2026.acl-tutorials)
Copied to clipboard
| Challenge: | This tutorial examines comprehensive evaluation strategies to assess the reasoning abilities of large language models (LLMs) advanced inference time methods and post-training methods that aim to make LLMs think more like humans are discussed in this tutorial. |
| Approach: | This tutorial explores comprehensive evaluation strategies to assess the reasoning abilities of large language models (LLMs) and discusses two types of methods to improve models’ reasoning: advanced inference time methods, structured and self-improvement inference methods, and post-training methods, such as RLHF, DPO, and GRPO. |
| Outcome: | This tutorial examines evaluation strategies to assess the reasoning abilities of large language models and discusses two types of methods to improve models’ reasoning. |