Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility (2025.acl-short)
Copied to clipboard
| Challenge: | Existing research on the linguistic capabilities of large language models has focused on their performance in language interpretation. |
| Approach: | They examine whether large language models (LLMs) process language similarly to humans . they use an empirically documented asymmetry between production and interpretation in humans a testbed . |
| Outcome: | The proposed model can replicate human-like distinctions between production and interpretation. |
Similar Papers
Large Language Models Are Partially Primed in Pronoun Interpretation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing studies suggest large language models acquire rich linguistic representations, but little is known about whether they adapt to linguistic biases in a human-like way. |
| Approach: | They examine whether large language models display human-like referential biases using stimuli and procedures from real psycholinguistic experiments. |
| Outcome: | The proposed models display human-like referential biases when exposed to referential patterns in the local context. |
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Human-LLM Dialogue (2026.findings-acl)
Copied to clipboard
Jonathan Ivey, Shivani Kumar, Jiayu Liu, Hua Shen, Sushrita Rakshit, Rohan Raju, Haotian Zhang, Aparna Ananthasubramaniam, Junghwan Kim, Bowen Yi, Dustin Wright, Abraham Israeli, Anders Giovanni Møller, Lechen Zhang, David Jurgens
| Challenge: | Recent work has sought to use large language models to simulate human-human and human-LLM interactions. |
| Approach: | They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts. |
| Outcome: | The proposed models perform similarly in simulating English, Chinese, and Russian dialogues. |
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models exhibit increasingly sophisticated linguistic capabilities, yet the extent to which these models reflect human-like cognition versus advanced pattern recognition remains an open question. |
| Approach: | They conduct a series of targeted experiments to assess whether LLMs construct semantic representations and pragmatic inferences in a human-like manner. |
| Outcome: | The proposed framework can be used to assess the cognitive and linguistic capabilities of large language models (LLMs). |
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly studied as repositories of linguistic knowledge. |
| Approach: | They compare LLMs’ performance as pragmatic listeners and as pragmatic speakers . they find a robust asymmetry between pragmatic evaluation and pragmatic generation . |
| Outcome: | The proposed models perform better as listeners than speakers, and produce more appropriate language than speakers. |
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects (2026.findings-acl)
Copied to clipboard
| Challenge: | a number of theories have been proposed to account for content effects in large language models, including the dual-process theory of reasoning, but the mechanisms behind content effects remain unclear. |
| Approach: | They propose to encode validity and plausibility concepts in LLMs by aligning them in representational geometry. |
| Outcome: | The proposed model conflates validity and plausibility, and vice versa. |
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have shown capabilities close to human performance in various analytical tasks. |
| Approach: | They investigate the efficiency and accuracy of Large Language Models in specialized tasks . they integrate LLMs with expert annotators to observe the impact of LLM suggestions . |
| Outcome: | The proposed model improves task completion speed but introduces anchoring bias . the proposed model is not suitable for open-ended analysis, but is capable of handling specialized tasks. |
Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives (2026.acl-long)
Copied to clipboard
| Challenge: | a new study examines whether large language models acquire embodied cognition and cultural conventions from training data . demonstratives are a natural lens for evaluating linguistic phenomena that reflect cultural variation . aaron e. duan and j. nà: "the complexity of the language model is a major challenge for LLMs" |
| Approach: | They introduce demonstratives as a probe for grounded knowledge by analyzing 6,400 responses from 320 native speakers. |
| Outcome: | The proposed model fails to understand proximal–distal contrast and shows no cultural differences . the proposed model is a new probe for evaluating embodied cognition and cultural conventions . |
Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on Understanding (2023.emnlp-main)
Copied to clipboard
| Challenge: | Current Large Language Models (LLMs) are unparalleled in their ability to generate grammatically correct, fluent text. |
| Approach: | They argue that LLMs only parrot statistical patterns in training data and that language learning in LLM cannot inform human language learning. |
| Outcome: | The proposed model can generate grammatically correct, fluent text without requiring human intervention. |
Large Language Models Help Humans Verify Truthfulness – Except When They Are Convincingly Wrong (2024.naacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used for accessing information on the web. |
| Approach: | They conduct experiments with 80 crowdworkers to compare LLMs with search engines . they ask LLM to provide contrastive information to reduce over-reliance on LLM . |
| Outcome: | The results show that LLMs can outperform search engines but not LLM explanations . the study shows that LMS explanations are not reliable replacements for reading retrieved passages compared to search engines alone. |
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models (2025.acl-long)
Copied to clipboard
| Challenge: | Modern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs’ and humans’ language processing. |
| Approach: | They propose to answer two questions: 1. What makes garden-path sentences hard for humans? 2. Do the same reasons make garden- path sentences hard? |
| Outcome: | The proposed models show that humans struggle with specific syntactic complexities, with some models showing high correlation with human comprehension. |