Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in LLMs have sparked a debate on whether they understand text. |
| Approach: | They propose two working definitions for understanding which explicitly acknowledge the question of consciousness and draw connections with a rich literature in philosophy, psychology and neuroscience. |
| Outcome: | The proposed models achieve impressive results on various benchmarks, seeming to generalize to unseen tasks and domains. |
Similar Papers
To Test Machine Comprehension, Start by Defining Comprehension (2020.acl-main)
Copied to clipboard
| Challenge: | Existing approaches to machine reading comprehension do not adequately define comprehension, authors argue . authors argue that existing systems are not up to the task of narrative understanding as they define it . |
| Approach: | They propose a detailed definition of comprehension for short narratives . they argue existing systems are not up to the task of narrative understanding . |
| Outcome: | The proposed task definitions suggest existing systems are not up to the task of narrative understanding as they define it. |
Do Language Models Have Semantics? On the Five Standard Positions (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are trained to solve the so-called cloze task . solving clozing tasks is essentially a memorization task, says a recent study . |
| Approach: | They propose to use five positions to determine whether large language models exhibit semantic understanding . large language model is trained to solve the so-called cloze task . |
| Outcome: | The proposed theory is based on a pairwise comparison of five positions on semantic understanding in large language models and chatbots. |
Social Intelligence in the Age of LLMs (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are a powerful tool for integrating human-like communication and context-aware interactions into artificial systems. |
| Approach: | They propose to introduce and overview different aspects of artificial social intelligence and their relationship with LLMs by introducing scientific methods for evaluating social intelligence in LLM. |
| Outcome: | This tutorial will introduce scientific methods for evaluating social intelligence in LLMs, highlighting the key challenges, and identifying promising research directions. |
Are NLP Models Good at Tracing Thoughts: An Overview of Narrative Understanding (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) excel in generating coherent texts, but their ability to comprehend the author’s thoughts remains uncertain. |
| Approach: | They conduct a comprehensive survey of narrative understanding tasks, examining their key features, definitions, taxonomy, associated datasets, evaluation metrics, and limitations. |
| Outcome: | The proposed framework could be extended to address novel narrative understanding tasks. |
Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data (2020.acl-main)
Copied to clipboard
| Challenge: | a priori, large neural language models are described as understanding or capturing meaning on tasks that are ostensibly meaningsensitive. |
| Approach: | They argue that a system trained only on form has no way to learn meaning . they argue that this is due to a misunderstanding of the relationship between form and meaning - which is a misconception in NLP . |
| Outcome: | The proposed model can't learn meaning because it only uses form as training data, the authors argue . they argue that a clear understanding of the distinction between form and meaning will guide the field towards better science around natural language understanding. |
Deep Bayesian Learning and Understanding (C18-3)
Copied to clipboard
| Challenge: | COLING 2018 is a conference for researchers and practitioners working on machine learning and deep learning. |
| Approach: | a tutorial on machine learning and deep learning will be presented at COLING 2018 . the tutorial will focus on statistical models, deep neural networks, sequential learning and natural language understanding . |
| Outcome: | This tutorial will present the latest advances in deep Bayesian and sequential learning at COLING 2018 . |
On General Language Understanding (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a recent paper suggests that the evidence underspecifies the understanding of large language models. |
| Approach: | They propose to use a "general language understanding" benchmark to examine what it could mean in machines. |
| Outcome: | The proposed model can be used to ground questions of the adequacy of benchmarking methods. |
A rebuttal of two common deflationary stances against LLM cognition (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are arguably the most predictive models of human cognition available. |
| Approach: | They argue that these deflationary claims need further justification . they argue that large language models are "just" simplistic entities . |
| Outcome: | The proposed models lack critical capacities, but they are not "just" models, the authors argue . they argue that the arguments need to be weighed against the evidence . |
Proceedings of the 2nd Workshop on Machine Reading for Question Answering (D19-58)
Copied to clipboard
| Challenge: | a workshop focuses on machine reading for question answering . despite recent progress, there is much to be desired about these datasets and systems . |
| Approach: | This year, they present a shared task on machine reading for question answering . they adapt and unified 18 distinct question answering datasets into the same format . |
| Outcome: | The proposed system achieves an average F1 score of 72.5 on the held-out datasets. |
The Thin Line Between Comprehension and Persuasion in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models are excellent at maintaining high-level, convincing dialogue . but it remains unclear whether their persuasive success reflects genuine understanding of the discourse . |
| Approach: | They examine whether LLMs' persuasive success reflects genuine understanding of the discourse . they find that LLM's effectively maintain coherent, persuasive debates . |
| Outcome: | The findings show that large language models can sway beliefs of participants and audiences. |