Challenge: Visual arguments rely on images to persuade viewers to do or believe something .
Approach: They propose three tasks for evaluating visual argument understanding . they use visual premises, commonsense premises and reasoning trees to analyze visual arguments .
Outcome: The proposed tasks evaluate visual argument understanding using a dataset of 1,611 images annotated with 5,112 visual premises (with regions), 5,574 commonsense premises, and reasoning trees connecting them into structured arguments.

Similar Papers

VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing Large Multi-modal Models lack a robust visual processing capability that is often masked by evaluation metrics that prioritize final-answer accuracy.
Approach: They propose a three-layer evaluation framework that scrutinizes the generation of valid visual aids and the soundness of subsequent reasoning steps.
Outcome: The proposed framework examines the generation of valid visual aids and the soundness of subsequent reasoning steps on state-of-the-art models.
Visualization of the Topic Space of Argument Search Results in args.me (D18-2)

Copied to clipboard

Challenge: args.me is the first search engine for controversial topics . it ranks pro and con arguments by their relevance to a topic .
Approach: They propose a visualization interface for result exploration that provides an overview of main aspects in a barycentric coordinate system.
Outcome: The proposed search engine is the first dedicated argument search engine on the web.
Visual Choice of Plausible Alternatives: An Evaluation of Image-based Commonsense Causal Reasoning (L18-1)

Copied to clipboard

Challenge: Existing methods for evaluating plausibility of events are focused on measuring causal dependency between events or actions.
Approach: They propose a task to identify the more plausible alternative with their commonsense causal context.
Outcome: The proposed task is based on a visual COPA dataset with 380 questions and over 1K images with various topics.
Looking Beyond the Pixels: Evaluating Visual Metaphor Understanding in VLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Visual metaphors are a complex vision–language phenomenon that requires both perceptual and conceptual reasoning to understand.
Approach: They introduce a visual metaphor dataset featuring 2177 synthetic and 350 human-annotated images and benchmark several SOTA VLMs on two tasks: Visual Metaphor Captioning (VMC) and Visual Metamorphosis VQA (VM-VQA).
Outcome: The proposed model outperforms standard few-shot baselines on visual metaphors and VM-VQA tasks.
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)

Copied to clipboard

Challenge: Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process.
Approach: They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions.
Outcome: The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input.
Natural Language Rationales with Full-Stack Visual Reasoning: From Pixels to Semantic Frames to Commonsense Graphs (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models that use natural language rationales provide intuitive, higher-level explanations that are easily understandable by humans.
Approach: They propose a model that generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs.
Outcome: The proposed model generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs.
Words Aren’t Enough, Their Order Matters: On the Robustness of Grounding Visual Referring Expressions (2020.acl-main)

Copied to clipboard

Challenge: Visual referring expression recognition is a task that requires natural language understanding in the context of an image.
Approach: They propose to use contrastive learning and multi-task learning to increase the robustness of ViLBERT, the current state-of-the-art model for this task.
Outcome: The proposed methods are 12% to 23% lower in performance than the established progress for this task.
Representing Verbs with Visual Argument Vectors (2020.lrec-1)

Copied to clipboard

Challenge: Existing models for verb semantic similarities are based on linguistic data, but they do not register intuitive attributes.
Approach: They evaluated two textual distributional semantic models and a visual one to explore verb semantic similarities.
Outcome: The proposed models extract meaningful information and capture semantic similarity between verbs using visual distributional models.
A negative case analysis of visual grounding methods for VQA (2020.acl-main)

Copied to clipboard

Challenge: Existing Visual Question Answering (VQA) methods exploit dataset biases and spurious statistical correlations instead of producing correct answers for the right reasons.
Approach: They propose to incorporate visual cues to better ground VQA models . they also propose a regularization effect which prevents over-fitting to linguistic priors .
Outcome: The proposed method outperforms existing methods on the Visual Question Answering (VQA) dataset.
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in language and vision assistants have showcased impressive capabilities but suffer from a lack of transparency, limiting broader research and reproducibility.
Approach: They propose to redefine the design of vision-language models by identifying key components and creating efficient models with constrained inference costs.
Outcome: The proposed models achieve significant improvements in inference throughput while maintaining high performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations