Challenge: Modern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs’ and humans’ language processing.
Approach: They propose to answer two questions: 1. What makes garden-path sentences hard for humans? 2. Do the same reasons make garden- path sentences hard?
Outcome: The proposed models show that humans struggle with specific syntactic complexities, with some models showing high correlation with human comprehension.

Similar Papers

Comparing human and language models sentence processing difficulties on complex structures (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) that converse with humans are a reality, but do LLMs experience human-like processing difficulties?
Approach: They systematically compare human and LLM sentence comprehension across seven challenging linguistic structures.
Outcome: The proposed model achieves near perfect accuracy on non-GP structures, but struggles on GP structures.
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead? (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models have shown capabilities close to human performance in various analytical tasks.
Approach: They investigate the efficiency and accuracy of Large Language Models in specialized tasks . they integrate LLMs with expert annotators to observe the impact of LLM suggestions .
Outcome: The proposed model improves task completion speed but introduces anchoring bias . the proposed model is not suitable for open-ended analysis, but is capable of handling specialized tasks.
Pun Unintended: LLMs and the Illusion of Humor Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models for pun detection lack nuanced grasp typical of human interpretation.
Approach: They analyze existing pun detection benchmarks and human evaluation across recent LLMs to find subtle changes in puns that mislead LLM.
Outcome: The proposed models lack the nuance typical of human interpretation and lack the depth of their analysis to detect puns.
Extracting structure from an LLM - how to improve on surprisal-based models of Human Language Processing (2025.coling-main)

Copied to clipboard

Challenge: Existing computational models capture prediction and reanalysis using Large Language Models (LLMs) and a statistical measure known as ‘surprisal’.
Approach: They propose to extract structural information from Large Language Models and a statistical measure known as ‘surprisal’ to integrate it with their learnt statistics.
Outcome: The proposed model achieved higher correlation with human reading times and better predicted the garden path effect and could distinguish between sentence types with different levels of difficulty.
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal (2026.acl-long)

Copied to clipboard

Challenge: Surprisal theory claims that difficulty of sentences increases linearly with surprise . a neural LM that can explain garden-path effects cannot be built, says a new study .
Approach: They propose to fine-tune neural LMs to better align surprisal-based reading-time estimates with actual reading times.
Outcome: a new study shows that fine-tuned neural LMs do not overfit on held-out items . the results show that they improve predictive power for human reading times .
Large Human Language Models: A Need and the Challenges (2024.naacl-long)

Copied to clipboard

Challenge: a growing recognition of the importance of modeling human and social factors into human-centered NLP models . authors advocate for three positions toward creating large human language models based on psychological and behavioral sciences .
Approach: et al. advocate for three positions toward creating large human language models . they argue that LM training should include the human context and recognize that people are more than their group .
Outcome: a new study shows that learning language from linguistic signals alone is not adequate, according to a recent paper . authors advocate for three positions toward creating large human language models . a human-centered model should include the human context, and account for the dynamic nature of the human environment, they say .
Large Language Models Are Partially Primed in Pronoun Interpretation (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies suggest large language models acquire rich linguistic representations, but little is known about whether they adapt to linguistic biases in a human-like way.
Approach: They examine whether large language models display human-like referential biases using stimuli and procedures from real psycholinguistic experiments.
Outcome: The proposed models display human-like referential biases when exposed to referential patterns in the local context.
LLMs meet Bloom’s Taxonomy: A Cognitive View on Large Language Model Evaluations (2025.coling-main)

Copied to clipboard

Challenge: Existing evaluation approaches for Large Language Models lack a structured approach that reflects the underlying cognitive abilities required for solving the tasks.
Approach: They propose a hierarchical approach to evaluation of Large Language Models that leverages Bloom’s Taxonomy to identify how well they cover the levels of Bloom’ s taxonomies.
Outcome: The proposed evaluation frameworks cover the Bloom’s Taxonomy, a hierarchical framework for categorizing cognitive skills, on the most widely used benchmarks.
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains.
Approach: They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks .
Outcome: The proposed evaluations are reproducible, reliable, and robust.
Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility (2025.acl-short)

Copied to clipboard

Challenge: Existing research on the linguistic capabilities of large language models has focused on their performance in language interpretation.
Approach: They examine whether large language models (LLMs) process language similarly to humans . they use an empirically documented asymmetry between production and interpretation in humans a testbed .
Outcome: The proposed model can replicate human-like distinctions between production and interpretation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations