Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan Acquisition (2024.acl-long)
Copied to clipboard
| Challenge: | Recent work on dialogue-based collaborative plan acquisition (CPA) suggests Theory of Mind (ToM) modelling can improve missing knowledge prediction in settings with asymmetric skill-sets and knowledge. |
| Approach: | They propose to use task-specific constraints to represent plans as graphs and exploit task-related constraints to improve missing knowledge prediction in CPA. |
| Outcome: | The proposed model improves missing knowledge prediction in contexts with asymmetric skill-sets and knowledge, but the improvements diminish . the proposed model is compared with baseline models and found to be more effective than existing models. |
Similar Papers
Beyond Words: Integrating Theory of Mind into Conversational Agents for Human-Like Belief, Desire, and Intention Alignment (2025.findings-acl)
Copied to clipboard
| Challenge: | Empirical evaluations of LLaMA-3 models demonstrate that ToM-informed alignment improves response quality, achieving win rates of 63% and 67%, respectively. |
| Approach: | They investigate whether open-source LLaMA models can represent and retain ToM-related constructs and whether they can be used to generate more aligned responses. |
| Outcome: | The proposed models can represent and retain ToM-related constructs and improve response quality. |
Theory of Mind in Large Language Models: Assessment and Enhancement (2025.acl-long)
Copied to clipboard
| Challenge: | Theory of Mind (ToM) is a cornerstone of human social intelligence . Large Language Models (LLMs) are increasingly integrated into daily life . |
| Approach: | They analyze evaluation benchmarks and enhancement strategies to evaluate LLMs' ToM capabilities. |
| Outcome: | The proposed and widely used story-based benchmarks and enhancement strategies are used to evaluate LLMs' ToM capabilities. |
Agentic-ToM: Cognition-Inspired Agentic Processing For Enhancing Theory of Mind Reasoning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Current models struggle with reasoning about others’ perspectives, limiting their ability to attribute mental states to oneself and others. |
| Approach: | They propose to embed psychologically-grounded functions into LLMs to enable them to attribute mental states to oneself and others, known as Theory of Mind. |
| Outcome: | The proposed approach outperforms baselines on three ToM datasets without task-specific modifications. |
Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition (2025.coling-main)
Copied to clipboard
| Challenge: | Theory of Mind (ToM) is the ability to attribute and infer the mental states of others. |
| Approach: | They propose an LLM-based inference algorithm that improves model performance on complex ToM tasks by simulating user perspectives. |
| Outcome: | The proposed algorithm improves model performance on complex ToM tasks while requiring minimal prompt tuning across tasks and no additional model training. |
ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing Theory of Mind (ToM) benchmarks focus on text-only or dyadic interactions, but to address this gap, we propose ToM-SSI: a new benchmark specifically designed to test ToM capabilities in environments rich with social interactions and spatial dynamics. |
| Approach: | They propose to use the Sally-Anne test to test ToM capabilities in environments rich in social interactions and spatial dynamics. |
| Outcome: | The proposed model captures a wider range of social cognition than existing models and demonstrates that existing models are still limited in these new tasks. |
Views Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground (2024.findings-acl)
Copied to clipboard
Adil Soubki, John Murzaku, Arash Yousefi Jordehi, Peter Zeng, Magdalena Markowska, Seyed Abolghasem Mirroshandel, Owen Rambow
| Challenge: | Existing benchmarks for theory of mind (ToM) use synthetic data, which can misalign with human behavior. |
| Approach: | They propose a question-answer benchmark based on naturally occurring spoken dialogs to evaluate theory of mind capabilities of language models. |
| Outcome: | The proposed dataset shows that LMs struggle to demonstrate theory of mind (ToM) . |
CoSToM: Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models lack intrinsic cognition and cannot generalize to complex task-specific scenarios. |
| Approach: | They propose a framework that transitions from mechanistic interpretation to active intervention to map internal distributions of ToM features and implement it via targeted activation steering within ToM-critical layers. |
| Outcome: | The proposed framework significantly enhances human-like social reasoning capabilities and dialogue quality. |
Infusing Theory of Mind into Socially Intelligent LLM Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | Theory of Mind (ToM) is a key aspect of human social intelligence, yet chatbots and LLMs do not typically integrate it. |
| Approach: | They propose a method that integrates Theory of Mind (ToM) into chatbots and dialogue agents to generate mental states between dialogue turns. |
| Outcome: | The proposed method improves dialogue and social interaction by integrating ToM with dialogue lookahead. |
Machine Theory of Mind Needs Machine Validation (2025.findings-acl)
Copied to clipboard
| Challenge: | In recent years there has been an explosion of interest in studying the extent to which language models (LMs) display a theory of mind (ToM) despite the growth of evaluation tools, the extent of evidence for ToM remains unclear. |
| Approach: | They conduct a survey of 16 recent studies aimed at measuring ToM in language models and found that only half do so for patterns only a machine might exploit. |
| Outcome: | The results show that the datasets that show high LM performance on ToM tasks are easier than their peers, likely due to the presence of spurious patterns in the data. |
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic nature of human-AI interactions. |
| Approach: | They propose a new paradigm of interactive ToM evaluation with both perspective and metric shifts. |
| Outcome: | The proposed approach improves the performance of four representative LLM enhancement techniques using real-world datasets and a user study. |