Challenge: Theory-of-Mind (ToM) is a psychological capability that allows humans to understand and interpret the mental states of others.
Approach: They propose a CharToM-QA benchmark to assess the importance of comprehensive contextual understanding about personal backgrounds in ToM.
Outcome: The proposed model outperforms existing models on 1,035 ToM questions based on classic novels and shows that educated participants perform better when they have read the novels than non-educated participants.

Similar Papers

Theory of Mind in Large Language Models: Assessment and Enhancement (2025.acl-long)

Copied to clipboard

Challenge: Theory of Mind (ToM) is a cornerstone of human social intelligence . Large Language Models (LLMs) are increasingly integrated into daily life .
Approach: They analyze evaluation benchmarks and enhancement strategies to evaluate LLMs' ToM capabilities.
Outcome: The proposed and widely used story-based benchmarks and enhancement strategies are used to evaluate LLMs' ToM capabilities.
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing N-ToM benchmarks lack ambiguous and artificial narratives, lack of personality traits and preferences, and limited diversity in the questions posed.
Approach: They propose a benchmark to assess Neural Theory-of-Mind (N-ToM) with longer and clearer narrative stories, characters with explicit personality traits, actions triggered by character intentions, and questions designed to challenge LLMs’ abilities of modeling characters’ mental states.
Outcome: The proposed test aims to assess the performance of LLMs in the physical and psychological worlds.
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for Theory of Mind (ToM) focus on whether agents have correct beliefs about others.
Approach: They propose to evaluate Theory of Mind (ToM) capabilities in Large Language Models (LLMs) they propose to use the theory of mind to determine whether and how to invoke ToM .
Outcome: The proposed frameworks can be used to evaluate the performance of large language models (LLMs) in biological agents.
Agentic-ToM: Cognition-Inspired Agentic Processing For Enhancing Theory of Mind Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: Current models struggle with reasoning about others’ perspectives, limiting their ability to attribute mental states to oneself and others.
Approach: They propose to embed psychologically-grounded functions into LLMs to enable them to attribute mental states to oneself and others, known as Theory of Mind.
Outcome: The proposed approach outperforms baselines on three ToM datasets without task-specific modifications.
Views Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground (2024.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for theory of mind (ToM) use synthetic data, which can misalign with human behavior.
Approach: They propose a question-answer benchmark based on naturally occurring spoken dialogs to evaluate theory of mind capabilities of language models.
Outcome: The proposed dataset shows that LMs struggle to demonstrate theory of mind (ToM) .
XToM: Exploring the Multilingual Theory of Mind for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of ToM in LLMs are limited to English, neglecting the linguistic diversity that shapes human cognition.
Approach: They propose a multilingual benchmark that evaluates ToM across five languages . they find that models excel in multilingual language understanding, but their ToM performance varies across languages.
Outcome: The proposed benchmark evaluates LLMs across five languages and incorporates diverse task scenarios.
Beyond Words: Integrating Theory of Mind into Conversational Agents for Human-Like Belief, Desire, and Intention Alignment (2025.findings-acl)

Copied to clipboard

Challenge: Empirical evaluations of LLaMA-3 models demonstrate that ToM-informed alignment improves response quality, achieving win rates of 63% and 67%, respectively.
Approach: They investigate whether open-source LLaMA models can represent and retain ToM-related constructs and whether they can be used to generate more aligned responses.
Outcome: The proposed models can represent and retain ToM-related constructs and improve response quality.
Machine Theory of Mind Needs Machine Validation (2025.findings-acl)

Copied to clipboard

Challenge: In recent years there has been an explosion of interest in studying the extent to which language models (LMs) display a theory of mind (ToM) despite the growth of evaluation tools, the extent of evidence for ToM remains unclear.
Approach: They conduct a survey of 16 recent studies aimed at measuring ToM in language models and found that only half do so for patterns only a machine might exploit.
Outcome: The results show that the datasets that show high LM performance on ToM tasks are easier than their peers, likely due to the presence of spurious patterns in the data.
Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that large language models (LLMs) reason about others' emotional states using contextual information, within a Theory-of-Mind framework.
Approach: They propose to use large language models to reason about others’ emotional states using contextual information within a Theory-of-Mind framework.
Outcome: The proposed models can reason about situations and appraisals, but are poor at associating situational outcomes and appraisal with specific emotions.
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic nature of human-AI interactions.
Approach: They propose a new paradigm of interactive ToM evaluation with both perspective and metric shifts.
Outcome: The proposed approach improves the performance of four representative LLM enhancement techniques using real-world datasets and a user study.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations