Challenge: Recent advances in LLMs have sparked a debate on whether they understand text.
Approach: They propose two working definitions for understanding which explicitly acknowledge the question of consciousness and draw connections with a rich literature in philosophy, psychology and neuroscience.
Outcome: The proposed models achieve impressive results on various benchmarks, seeming to generalize to unseen tasks and domains.

Similar Papers

To Test Machine Comprehension, Start by Defining Comprehension (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to machine reading comprehension do not adequately define comprehension, authors argue . authors argue that existing systems are not up to the task of narrative understanding as they define it .
Approach: They propose a detailed definition of comprehension for short narratives . they argue existing systems are not up to the task of narrative understanding .
Outcome: The proposed task definitions suggest existing systems are not up to the task of narrative understanding as they define it.
Do Language Models Have Semantics? On the Five Standard Positions (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are trained to solve the so-called cloze task . solving clozing tasks is essentially a memorization task, says a recent study .
Approach: They propose to use five positions to determine whether large language models exhibit semantic understanding . large language model is trained to solve the so-called cloze task .
Outcome: The proposed theory is based on a pairwise comparison of five positions on semantic understanding in large language models and chatbots.
Social Intelligence in the Age of LLMs (2025.naacl-tutorial)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a powerful tool for integrating human-like communication and context-aware interactions into artificial systems.
Approach: They propose to introduce and overview different aspects of artificial social intelligence and their relationship with LLMs by introducing scientific methods for evaluating social intelligence in LLM.
Outcome: This tutorial will introduce scientific methods for evaluating social intelligence in LLMs, highlighting the key challenges, and identifying promising research directions.
Are NLP Models Good at Tracing Thoughts: An Overview of Narrative Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) excel in generating coherent texts, but their ability to comprehend the author’s thoughts remains uncertain.
Approach: They conduct a comprehensive survey of narrative understanding tasks, examining their key features, definitions, taxonomy, associated datasets, evaluation metrics, and limitations.
Outcome: The proposed framework could be extended to address novel narrative understanding tasks.
Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data (2020.acl-main)

Copied to clipboard

Challenge: a priori, large neural language models are described as understanding or capturing meaning on tasks that are ostensibly meaningsensitive.
Approach: They argue that a system trained only on form has no way to learn meaning . they argue that this is due to a misunderstanding of the relationship between form and meaning - which is a misconception in NLP .
Outcome: The proposed model can't learn meaning because it only uses form as training data, the authors argue . they argue that a clear understanding of the distinction between form and meaning will guide the field towards better science around natural language understanding.
Deep Bayesian Learning and Understanding (C18-3)

Copied to clipboard

Challenge: COLING 2018 is a conference for researchers and practitioners working on machine learning and deep learning.
Approach: a tutorial on machine learning and deep learning will be presented at COLING 2018 . the tutorial will focus on statistical models, deep neural networks, sequential learning and natural language understanding .
Outcome: This tutorial will present the latest advances in deep Bayesian and sequential learning at COLING 2018 .
On General Language Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: a recent paper suggests that the evidence underspecifies the understanding of large language models.
Approach: They propose to use a "general language understanding" benchmark to examine what it could mean in machines.
Outcome: The proposed model can be used to ground questions of the adequacy of benchmarking methods.
A rebuttal of two common deflationary stances against LLM cognition (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are arguably the most predictive models of human cognition available.
Approach: They argue that these deflationary claims need further justification . they argue that large language models are "just" simplistic entities .
Outcome: The proposed models lack critical capacities, but they are not "just" models, the authors argue . they argue that the arguments need to be weighed against the evidence .
Proceedings of the 2nd Workshop on Machine Reading for Question Answering (D19-58)

Copied to clipboard

Challenge: a workshop focuses on machine reading for question answering . despite recent progress, there is much to be desired about these datasets and systems .
Approach: This year, they present a shared task on machine reading for question answering . they adapt and unified 18 distinct question answering datasets into the same format .
Outcome: The proposed system achieves an average F1 score of 72.5 on the held-out datasets.
The Thin Line Between Comprehension and Persuasion in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Large language models are excellent at maintaining high-level, convincing dialogue . but it remains unclear whether their persuasive success reflects genuine understanding of the discourse .
Approach: They examine whether LLMs' persuasive success reflects genuine understanding of the discourse . they find that LLM's effectively maintain coherent, persuasive debates .
Outcome: The findings show that large language models can sway beliefs of participants and audiences.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations