Challenge: Recent work on dialogue-based collaborative plan acquisition (CPA) suggests Theory of Mind (ToM) modelling can improve missing knowledge prediction in settings with asymmetric skill-sets and knowledge.
Approach: They propose to use task-specific constraints to represent plans as graphs and exploit task-related constraints to improve missing knowledge prediction in CPA.
Outcome: The proposed model improves missing knowledge prediction in contexts with asymmetric skill-sets and knowledge, but the improvements diminish . the proposed model is compared with baseline models and found to be more effective than existing models.

Similar Papers

Beyond Words: Integrating Theory of Mind into Conversational Agents for Human-Like Belief, Desire, and Intention Alignment (2025.findings-acl)

Copied to clipboard

Challenge: Empirical evaluations of LLaMA-3 models demonstrate that ToM-informed alignment improves response quality, achieving win rates of 63% and 67%, respectively.
Approach: They investigate whether open-source LLaMA models can represent and retain ToM-related constructs and whether they can be used to generate more aligned responses.
Outcome: The proposed models can represent and retain ToM-related constructs and improve response quality.
Theory of Mind in Large Language Models: Assessment and Enhancement (2025.acl-long)

Copied to clipboard

Challenge: Theory of Mind (ToM) is a cornerstone of human social intelligence . Large Language Models (LLMs) are increasingly integrated into daily life .
Approach: They analyze evaluation benchmarks and enhancement strategies to evaluate LLMs' ToM capabilities.
Outcome: The proposed and widely used story-based benchmarks and enhancement strategies are used to evaluate LLMs' ToM capabilities.
Agentic-ToM: Cognition-Inspired Agentic Processing For Enhancing Theory of Mind Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: Current models struggle with reasoning about others’ perspectives, limiting their ability to attribute mental states to oneself and others.
Approach: They propose to embed psychologically-grounded functions into LLMs to enable them to attribute mental states to oneself and others, known as Theory of Mind.
Outcome: The proposed approach outperforms baselines on three ToM datasets without task-specific modifications.
Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition (2025.coling-main)

Copied to clipboard

Challenge: Theory of Mind (ToM) is the ability to attribute and infer the mental states of others.
Approach: They propose an LLM-based inference algorithm that improves model performance on complex ToM tasks by simulating user perspectives.
Outcome: The proposed algorithm improves model performance on complex ToM tasks while requiring minimal prompt tuning across tasks and no additional model training.
ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions (2025.emnlp-main)

Copied to clipboard

Challenge: Existing Theory of Mind (ToM) benchmarks focus on text-only or dyadic interactions, but to address this gap, we propose ToM-SSI: a new benchmark specifically designed to test ToM capabilities in environments rich with social interactions and spatial dynamics.
Approach: They propose to use the Sally-Anne test to test ToM capabilities in environments rich in social interactions and spatial dynamics.
Outcome: The proposed model captures a wider range of social cognition than existing models and demonstrates that existing models are still limited in these new tasks.
Views Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground (2024.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for theory of mind (ToM) use synthetic data, which can misalign with human behavior.
Approach: They propose a question-answer benchmark based on naturally occurring spoken dialogs to evaluate theory of mind capabilities of language models.
Outcome: The proposed dataset shows that LMs struggle to demonstrate theory of mind (ToM) .
CoSToM: Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models lack intrinsic cognition and cannot generalize to complex task-specific scenarios.
Approach: They propose a framework that transitions from mechanistic interpretation to active intervention to map internal distributions of ToM features and implement it via targeted activation steering within ToM-critical layers.
Outcome: The proposed framework significantly enhances human-like social reasoning capabilities and dialogue quality.
Infusing Theory of Mind into Socially Intelligent LLM Agents (2026.findings-acl)

Copied to clipboard

Challenge: Theory of Mind (ToM) is a key aspect of human social intelligence, yet chatbots and LLMs do not typically integrate it.
Approach: They propose a method that integrates Theory of Mind (ToM) into chatbots and dialogue agents to generate mental states between dialogue turns.
Outcome: The proposed method improves dialogue and social interaction by integrating ToM with dialogue lookahead.
Machine Theory of Mind Needs Machine Validation (2025.findings-acl)

Copied to clipboard

Challenge: In recent years there has been an explosion of interest in studying the extent to which language models (LMs) display a theory of mind (ToM) despite the growth of evaluation tools, the extent of evidence for ToM remains unclear.
Approach: They conduct a survey of 16 recent studies aimed at measuring ToM in language models and found that only half do so for patterns only a machine might exploit.
Outcome: The results show that the datasets that show high LM performance on ToM tasks are easier than their peers, likely due to the presence of spurious patterns in the data.
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic nature of human-AI interactions.
Approach: They propose a new paradigm of interactive ToM evaluation with both perspective and metric shifts.
Outcome: The proposed approach improves the performance of four representative LLM enhancement techniques using real-world datasets and a user study.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations