Challenge: Existing benchmarks for theory of mind are flawed due to dataset biases . evaluators have been using the Sally-Anne test to infer false beliefs in others .
Approach: They propose to use question answering to evaluate theory of mind . they propose to explicitly control for data regularities via a careful examination of the answer space .
Outcome: The proposed evaluation protocol and dataset control for data regularities via a careful examination of the answer space.

Similar Papers

Evaluating Theory of Mind in Question Answering (D18-1)

Copied to clipboard

Challenge: a dataset is proposed for question answering models with respect to their capacity to reason about beliefs.
Approach: They propose a dataset for evaluating question answering models with respect to their capacity to reason about beliefs.
Outcome: The proposed dataset is inspired by theory-of-mind experiments that examine whether children are able to reason about beliefs of others.
Evaluation Paradigms in Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Despite substantial overlap, subtle but significant distinctions exert an outsize influence on research . one paradigm values creating more intelligent QA systems, the other paradigm values building QA system that appeals to users.
Approach: They propose to use the Cranfield and Manchester paradigms to describe research working towards building human-like, intelligent QA systems.
Outcome: The proposed paradigms are based on the findings of two recent studies on question answering (QA) the Cranfield paradigm is not new, but the Manchester paradigm is christened as the most eclectic in QA .
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic nature of human-AI interactions.
Approach: They propose a new paradigm of interactive ToM evaluation with both perspective and metric shifts.
Outcome: The proposed approach improves the performance of four representative LLM enhancement techniques using real-world datasets and a user study.
Model Analysis & Evaluation for Ambiguous Question Answering (2023.findings-acl)

Copied to clipboard

Challenge: Ambiguous questions are a challenge for Question Answering models as they require answers that cover multiple interpretations of the original query.
Approach: They aim to investigate whether model/data scaling improves the answers’ quality and whether automated metrics align with human judgment.
Outcome: The proposed models can generate long-form answers that combine conflicting information and provide valuable insights into the limitations of the current approaches.
Are Red Roses Red? Evaluating Consistency of Question-Answering Models (P19-1)

Copied to clipboard

Challenge: Existing question-answering systems are limited in their ability to test reasoning and comprehension.
Approach: They propose a method to automatically extract implications from QA datasets to evaluate models' consistency . they use a heuristic to generate such questions and retrain models with implication-augmented data .
Outcome: The proposed method shows that generated implications are well formed and valid . retraining with implication-augmented data improves consistency on both synthetic and human-generated implications.
Theory of Mind in Large Language Models: Assessment and Enhancement (2025.acl-long)

Copied to clipboard

Challenge: Theory of Mind (ToM) is a cornerstone of human social intelligence . Large Language Models (LLMs) are increasingly integrated into daily life .
Approach: They analyze evaluation benchmarks and enhancement strategies to evaluate LLMs' ToM capabilities.
Outcome: The proposed and widely used story-based benchmarks and enhancement strategies are used to evaluate LLMs' ToM capabilities.
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for Theory of Mind (ToM) focus on whether agents have correct beliefs about others.
Approach: They propose to evaluate Theory of Mind (ToM) capabilities in Large Language Models (LLMs) they propose to use the theory of mind to determine whether and how to invoke ToM .
Outcome: The proposed frameworks can be used to evaluate the performance of large language models (LLMs) in biological agents.
A Review on Deep Learning Techniques Applied to Answer Selection (C18-1)

Copied to clipboard

Challenge: Existing deep learning methods for answer selection are not feature engineering or expensive external resources.
Approach: They propose to use deep learning methods to analyze and predict answer quality . they use a set of candidate answers to identify which of the candidates answers the question correctly.
Outcome: The proposed methods produce impressive performance without feature engineering or expensive external resources.
Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Table Question Answering (TQA) aims to answer natural language questions using tabular data.
Approach: They propose a systematic overview of TQA research using large language models and summarize available benchmarks based on task features.
Outcome: The proposed framework provides a comprehensive overview of the current state of the art in the field of Table Question Answering.
On the Interaction of Belief Bias and Explanations (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to evaluate explainability fail to account for belief biases affecting human performance . previous studies have shown that neural models can make confident predictions relying on artifacts .
Approach: They propose to account for belief bias in explainability by using models of varying quality and adversarial examples.
Outcome: The proposed methods show that results change when using models of varying quality and adversarial examples.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations