Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding (2024.emnlp-main)
Copied to clipboard
| Challenge: | Visual arguments rely on images to persuade viewers to do or believe something . |
| Approach: | They propose three tasks for evaluating visual argument understanding . they use visual premises, commonsense premises and reasoning trees to analyze visual arguments . |
| Outcome: | The proposed tasks evaluate visual argument understanding using a dataset of 1,611 images annotated with 5,112 visual premises (with regions), 5,574 commonsense premises, and reasoning trees connecting them into structured arguments. |
Similar Papers
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing Large Multi-modal Models lack a robust visual processing capability that is often masked by evaluation metrics that prioritize final-answer accuracy. |
| Approach: | They propose a three-layer evaluation framework that scrutinizes the generation of valid visual aids and the soundness of subsequent reasoning steps. |
| Outcome: | The proposed framework examines the generation of valid visual aids and the soundness of subsequent reasoning steps on state-of-the-art models. |
Visualization of the Topic Space of Argument Search Results in args.me (D18-2)
Copied to clipboard
Yamen Ajjour, Henning Wachsmuth, Dora Kiesel, Patrick Riehmann, Fan Fan, Giuliano Castiglia, Rosemary Adejoh, Bernd Fröhlich, Benno Stein
| Challenge: | args.me is the first search engine for controversial topics . it ranks pro and con arguments by their relevance to a topic . |
| Approach: | They propose a visualization interface for result exploration that provides an overview of main aspects in a barycentric coordinate system. |
| Outcome: | The proposed search engine is the first dedicated argument search engine on the web. |
Visual Choice of Plausible Alternatives: An Evaluation of Image-based Commonsense Causal Reasoning (L18-1)
Copied to clipboard
Jinyoung Yeo, Gyeongbok Lee, Gengyu Wang, Seungtaek Choi, Hyunsouk Cho, Reinald Kim Amplayo, Seung-won Hwang
| Challenge: | Existing methods for evaluating plausibility of events are focused on measuring causal dependency between events or actions. |
| Approach: | They propose a task to identify the more plausible alternative with their commonsense causal context. |
| Outcome: | The proposed task is based on a visual COPA dataset with 380 questions and over 1K images with various topics. |
Looking Beyond the Pixels: Evaluating Visual Metaphor Understanding in VLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Visual metaphors are a complex vision–language phenomenon that requires both perceptual and conceptual reasoning to understand. |
| Approach: | They introduce a visual metaphor dataset featuring 2177 synthetic and 350 human-annotated images and benchmark several SOTA VLMs on two tasks: Visual Metaphor Captioning (VMC) and Visual Metamorphosis VQA (VM-VQA). |
| Outcome: | The proposed model outperforms standard few-shot baselines on visual metaphors and VM-VQA tasks. |
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)
Copied to clipboard
| Challenge: | Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process. |
| Approach: | They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions. |
| Outcome: | The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input. |
Natural Language Rationales with Full-Stack Visual Reasoning: From Pixels to Semantic Frames to Commonsense Graphs (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models that use natural language rationales provide intuitive, higher-level explanations that are easily understandable by humans. |
| Approach: | They propose a model that generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs. |
| Outcome: | The proposed model generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs. |
Words Aren’t Enough, Their Order Matters: On the Robustness of Grounding Visual Referring Expressions (2020.acl-main)
Copied to clipboard
| Challenge: | Visual referring expression recognition is a task that requires natural language understanding in the context of an image. |
| Approach: | They propose to use contrastive learning and multi-task learning to increase the robustness of ViLBERT, the current state-of-the-art model for this task. |
| Outcome: | The proposed methods are 12% to 23% lower in performance than the established progress for this task. |
Representing Verbs with Visual Argument Vectors (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing models for verb semantic similarities are based on linguistic data, but they do not register intuitive attributes. |
| Approach: | They evaluated two textual distributional semantic models and a visual one to explore verb semantic similarities. |
| Outcome: | The proposed models extract meaningful information and capture semantic similarity between verbs using visual distributional models. |
A negative case analysis of visual grounding methods for VQA (2020.acl-main)
Copied to clipboard
| Challenge: | Existing Visual Question Answering (VQA) methods exploit dataset biases and spurious statistical correlations instead of producing correct answers for the right reasons. |
| Approach: | They propose to incorporate visual cues to better ground VQA models . they also propose a regularization effect which prevents over-fitting to linguistic priors . |
| Outcome: | The proposed method outperforms existing methods on the Visual Question Answering (VQA) dataset. |
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in language and vision assistants have showcased impressive capabilities but suffer from a lack of transparency, limiting broader research and reproducibility. |
| Approach: | They propose to redefine the design of vision-language models by identifying key components and creating efficient models with constrained inference costs. |
| Outcome: | The proposed models achieve significant improvements in inference throughput while maintaining high performance. |