| Challenge: | Social IQa contains 38,000 multiple choice questions for probing emotional and social intelligence in a variety of everyday situations. |
| Approach: | They propose a crowdsourcing framework that collects commonsense questions along with correct and incorrect answers about social interactions. |
| Outcome: | The proposed framework mitigates stylistic artifacts in incorrect answers by asking workers to provide the right answer to a different but related question. |
Similar Papers
DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence Understanding (2023.emnlp-main)
Copied to clipboard
| Challenge: | Social intelligence is essential for understanding and reasoning about human expressions, intents and interactions. |
| Approach: | They propose a methodology to study the soundness of Social-IQ by applying simple perturbations to a dataset of multiple choice questions on videos of complex social interactions. |
| Outcome: | The proposed method reduces biases in the original dataset and improves performance. |
Commonsense Reasoning for Natural Language Processing (2020.acl-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we will outline the various types of commonsense knowledge and discuss techniques to gather and represent commonsence knowledge. |
| Approach: | This tutorial will provide researchers with the critical foundations and recent advances in commonsense representation and reasoning. |
| Outcome: | This tutorial will outline the various types of commonsense and discuss techniques to gather and represent commonsence knowledge while highlighting the challenges specific to this type of knowledge (e.g., reporting bias). |
PragmatiCQA: A Dataset for Pragmatic Question Answering in Conversations (2023.findings-acl)
Copied to clipboard
| Challenge: | Mars? - PragmatiCQA |
| Approach: | Mars? - The Paper . |
| Outcome: | The proposed dataset features 6873 QA pairs that explores pragmatic reasoning in conversations over a diverse set of topics. |
Towards Generalizable Neuro-Symbolic Systems for Commonsense Question Answering (D19-60)
Copied to clipboard
| Challenge: | Recent approaches on non-extractive commonsense QA show increased performance . attention-based injection seems to be preferable for knowledge integration . |
| Approach: | They propose to use attention-based injection to integrate knowledge into commonsense QA models. |
| Outcome: | The proposed methods show that attention-based injection is preferable for knowledge integration, and that the degree of domain overlap plays a crucial role in determining model success. |
Social Intelligence in the Age of LLMs (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are a powerful tool for integrating human-like communication and context-aware interactions into artificial systems. |
| Approach: | They propose to introduce and overview different aspects of artificial social intelligence and their relationship with LLMs by introducing scientific methods for evaluating social intelligence in LLM. |
| Outcome: | This tutorial will introduce scientific methods for evaluating social intelligence in LLMs, highlighting the key challenges, and identifying promising research directions. |
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge (N19-1)
Copied to clipboard
| Challenge: | Recent work on question answering relies on factoid questions with little general knowledge. |
| Approach: | They propose a dataset to capture commonsense question answering with prior knowledge . they extract multiple-choice questions that discriminate between the source and target concepts . |
| Outcome: | The proposed dataset captures commonsense reasoning beyond associations . it obtains 56% accuracy, well below human performance, which is 89% . |
Social Commonsense for Explanation and Cultural Bias Discovery (2023.eacl-main)
Copied to clipboard
| Challenge: | Social commonsense contains many human biases due to social and cultural influence. |
| Approach: | They aim to identify cultural biases in data that strongly influence model decisions . they use social commonsense knowledge to augment large-scale language models . |
| Outcome: | The proposed method shows that social commonsense knowledge can explain model behavior on two social tasks. |
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It (2026.findings-eacl)
Copied to clipboard
| Challenge: | Using social reasoning benchmarks, we uncover pervasive flaws in both benchmark items and evaluation methodology. |
| Approach: | They audit three widely used social reasoning benchmarks and identify flaws in their design and evaluation methodology. |
| Outcome: | The results challenge the validity of current benchmark-based claims about social reasoning in large language models. |
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios (2025.naacl-long)
Copied to clipboard
Xinyi Mou, Jingcong Liang, Jiayu Lin, Xinnong Zhang, Xiawei Liu, Shiyue Yang, Rong Ye, Lei Chen, Haoyu Kuang, Xuanjing Huang, Zhongyu Wei
| Challenge: | Large language models are increasingly employed to empower autonomous agents to simulate human behavior. |
| Approach: | They propose to evaluate LLM-driven agents through multi-turn interactions using a bottom-up approach to create diverse social scenarios constructed from extensive scripts. |
| Outcome: | The proposed model evaluates LLM-driven agents through multi-turn interactions emphasizing goal completion and implicit reasoning. |
Event2Mind: Commonsense Inference on Events, Intents, and Reactions (P18-1)
Copied to clipboard
| Challenge: | Using a crowdsourced corpus of 25,000 event phrases, we construct a new task that uses commonsense reasoning to reason about the likely intents and reactions of the event participants. |
| Approach: | They construct a crowdsourced corpus of 25,000 event phrases and use them to construct 'commonsense inference' they demonstrate that neural encoder-decoder models can compose embedding representations of previously unseen events and reason about the likely intents and reactions of the event participants. |
| Outcome: | The proposed task can be used to uncover implicit gender inequality in movie scripts. |