V-ALPHASOCIAL: Benchmark and Self-Reflective Chain-of-Thought Generation for Visual Social Commonsense Reasoning (2025.findings-acl)
Copied to clipboard
Zongyu Lin, Zhikun Xu, Xiaohan Song, Yixin Wan, Xingcheng Yao, Tsung-Han Lin, Selina Song, Pranav Subbaraman, Ben Zhou, Kai-Wei Chang, Yizhou Sun
| Challenge: | Social commonsense reasoning is a multimodal task that requires both textual and visual cues. |
| Approach: | They propose a method that integrates visual cues into social commonsense reasoning tasks. |
| Outcome: | The proposed method improves social commonsense reasoning on a multimodal foundation model. |
Similar Papers
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models (2024.naacl-long)
Copied to clipboard
| Challenge: | Vision-language models have demonstrated strong efficacy as visual assistants . however, evaluation of their reasoning capabilities requires a costly benchmark . |
| Approach: | They propose a pipeline to measure the reasoning consistency of vision-language models . they propose supervised fine-tuning of VLMs and feedback from LLMs . |
| Outcome: | The proposed framework reduces cost while ensuring the generation of a high-quality dataset. |
VIBE: Can a VLM Read the Room? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Vision Language Models (LLMs) cannot account for the role that non-verbal cues play in understanding social situations. |
| Approach: | They propose a task to test the capabilities of Vision Language Models (VLMs) to account for the visual social-pragmatic inference gap. |
| Outcome: | The proposed task tests the capabilities of a VLM for a social reasoning task. |
VChain: Chain-of-Visual-Thought for Reasoning in Video Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent video generation models struggle to synthesize complex dynamics with a coherent chain of consequences. |
| Approach: | They propose a framework that injects visual reasoning signals from multimodal models into video generation. |
| Outcome: | a new framework that leverages multimodal models to generate sparse keyframes significantly improves quality of generated videos. |
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for Multimodal Large Language Models (MLLMs) have been lacking due to the rich nature of social interaction. |
| Approach: | They propose a video benchmark to evaluate MLLMs' capabilities across social scene understanding, social state reasoning, and social dynamics prediction. |
| Outcome: | The proposed benchmarks evaluate MLLMs' capabilities across social scene understanding, social state reasoning, and social dynamics prediction tasks. |
Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference Tuning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to assess social-pragmatic inference in large language models are inadequacy, and preferential tuning is the best approach. |
| Approach: | They propose to use free-form models' responses as a measure to assess social-pragmatic reasoning and advocate for preference optimization over supervised finetuning (SFT). |
| Outcome: | The proposed model outperforms supervised finetuning (SFT) and offers a near-free launch in pragmatic abilities without compromising general capabilities. |
Does Chain-of-Thought Reasoning Help Mobile GUI Agents? An Empirical Study (2026.findings-acl)
Copied to clipboard
| Challenge: | Reasoning capabilities have improved vision-language models in domains like math, coding, and visual question-answering, but their impact on real-world applications remains unclear. |
| Approach: | They evaluate six pairs of VLMs by comparing their base and reasoning-enhanced versions across static and interactive benchmarks. |
| Outcome: | The reasoning-enhanced models perform better on static and interactive benchmarks than non-reasoning models. |
Beyond Semantic Similarity: Appraisal-Guided Chain-of-Thought Reasoning and Retrieval for Multimodal Emotional Support Conversations (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing retrieval-augmented generation paradigms rely on semantic similarity to retrieve historical dialogues that are surface analogous but therapeutically incongruent. |
| Approach: | They propose to use appraisal-guided reasoning chains to generate appraisal-based reasoning chains and apply a dual-signal verification mechanism to verify and correct them. |
| Outcome: | Extensive experiments on two ESC benchmarks show that the proposed model significantly outperforms state-of-the-art models. |
Sibyl: Empowering Empathetic Dialogue Generation in Large Language Models via Sensible and Visionary Commonsense Inference (2025.coling-main)
Copied to clipboard
Lanrui Wang, Jiangnan Li, Chenxu Yang, Zheng Lin, Hongyin Tang, Huan Liu, Yanan Cao, Jingang Wang, Weiping Wang
| Challenge: | Recent studies have focused on integrating commonsense knowledge into chatbots to enhance their ability to understand and generate dialogue responses. |
| Approach: | They propose a framework that integrates commonsense knowledge into chatbots to enable them to elicit more empathetic responses. |
| Outcome: | The proposed framework enables LLMs to uncover the implicit requirements of the conversation, aiming to elicit more empathetic responses. |
Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing models that understand image and text but also cross-reference in-between are lacking in evaluation data resources. |
| Approach: | They propose a multimodal evaluation pipeline to automatically generate question-answer pairs to test models’ understanding of the visual scene, text, and related knowledge. |
| Outcome: | The proposed model can answer the highly semantic VCR question correctly but fails to answer related visual question (Q2), textual question (q3), and background knowledge question ( Q4) as shallow mappings with language priors and unbalanced utilization of information between modalities. |
Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing evaluations for visual hallucinations are narrow. |
| Approach: | They propose a framework that decomposes reasoning chains into perception versus reasoning steps and uses off-the-shelf VLM judges for step-level faithfulness. |
| Outcome: | The proposed framework reduces Unfaithful Perception Rate while preserving final-answer accuracy. |