When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models (2025.acl-long)
Copied to clipboard
| Challenge: | Modern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs’ and humans’ language processing. |
| Approach: | They propose to answer two questions: 1. What makes garden-path sentences hard for humans? 2. Do the same reasons make garden- path sentences hard? |
| Outcome: | The proposed models show that humans struggle with specific syntactic complexities, with some models showing high correlation with human comprehension. |
Similar Papers
Comparing human and language models sentence processing difficulties on complex structures (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) that converse with humans are a reality, but do LLMs experience human-like processing difficulties? |
| Approach: | They systematically compare human and LLM sentence comprehension across seven challenging linguistic structures. |
| Outcome: | The proposed model achieves near perfect accuracy on non-GP structures, but struggles on GP structures. |
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have shown capabilities close to human performance in various analytical tasks. |
| Approach: | They investigate the efficiency and accuracy of Large Language Models in specialized tasks . they integrate LLMs with expert annotators to observe the impact of LLM suggestions . |
| Outcome: | The proposed model improves task completion speed but introduces anchoring bias . the proposed model is not suitable for open-ended analysis, but is capable of handling specialized tasks. |
Pun Unintended: LLMs and the Illusion of Humor Understanding (2025.emnlp-main)
Copied to clipboard
Alessandro Zangari, Matteo Marcuzzo, Andrea Albarelli, Mohammad Taher Pilehvar, Jose Camacho-Collados
| Challenge: | Existing models for pun detection lack nuanced grasp typical of human interpretation. |
| Approach: | They analyze existing pun detection benchmarks and human evaluation across recent LLMs to find subtle changes in puns that mislead LLM. |
| Outcome: | The proposed models lack the nuance typical of human interpretation and lack the depth of their analysis to detect puns. |
Extracting structure from an LLM - how to improve on surprisal-based models of Human Language Processing (2025.coling-main)
Copied to clipboard
| Challenge: | Existing computational models capture prediction and reanalysis using Large Language Models (LLMs) and a statistical measure known as ‘surprisal’. |
| Approach: | They propose to extract structural information from Large Language Models and a statistical measure known as ‘surprisal’ to integrate it with their learnt statistics. |
| Outcome: | The proposed model achieved higher correlation with human reading times and better predicted the garden path effect and could distinguish between sentence types with different levels of difficulty. |
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal (2026.acl-long)
Copied to clipboard
| Challenge: | Surprisal theory claims that difficulty of sentences increases linearly with surprise . a neural LM that can explain garden-path effects cannot be built, says a new study . |
| Approach: | They propose to fine-tune neural LMs to better align surprisal-based reading-time estimates with actual reading times. |
| Outcome: | a new study shows that fine-tuned neural LMs do not overfit on held-out items . the results show that they improve predictive power for human reading times . |
Large Human Language Models: A Need and the Challenges (2024.naacl-long)
Copied to clipboard
| Challenge: | a growing recognition of the importance of modeling human and social factors into human-centered NLP models . authors advocate for three positions toward creating large human language models based on psychological and behavioral sciences . |
| Approach: | et al. advocate for three positions toward creating large human language models . they argue that LM training should include the human context and recognize that people are more than their group . |
| Outcome: | a new study shows that learning language from linguistic signals alone is not adequate, according to a recent paper . authors advocate for three positions toward creating large human language models . a human-centered model should include the human context, and account for the dynamic nature of the human environment, they say . |
Large Language Models Are Partially Primed in Pronoun Interpretation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing studies suggest large language models acquire rich linguistic representations, but little is known about whether they adapt to linguistic biases in a human-like way. |
| Approach: | They examine whether large language models display human-like referential biases using stimuli and procedures from real psycholinguistic experiments. |
| Outcome: | The proposed models display human-like referential biases when exposed to referential patterns in the local context. |
LLMs meet Bloom’s Taxonomy: A Cognitive View on Large Language Model Evaluations (2025.coling-main)
Copied to clipboard
| Challenge: | Existing evaluation approaches for Large Language Models lack a structured approach that reflects the underlying cognitive abilities required for solving the tasks. |
| Approach: | They propose a hierarchical approach to evaluation of Large Language Models that leverages Bloom’s Taxonomy to identify how well they cover the levels of Bloom’ s taxonomies. |
| Outcome: | The proposed evaluation frameworks cover the Bloom’s Taxonomy, a hierarchical framework for categorizing cognitive skills, on the most widely used benchmarks. |
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)
Copied to clipboard
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty, Jimmy Huang
| Challenge: | Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains. |
| Approach: | They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks . |
| Outcome: | The proposed evaluations are reproducible, reliable, and robust. |
Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility (2025.acl-short)
Copied to clipboard
| Challenge: | Existing research on the linguistic capabilities of large language models has focused on their performance in language interpretation. |
| Approach: | They examine whether large language models (LLMs) process language similarly to humans . they use an empirically documented asymmetry between production and interpretation in humans a testbed . |
| Outcome: | The proposed model can replicate human-like distinctions between production and interpretation. |