Challenge: Recent advances in Instruction-fine-tuned Vision and Language Models (IVLMs) have prompted some studies to analyze the reasoning capabilities of IVLMs.
Approach: They introduce a vision and language task for Inductive Visual Reasoning that uses common attributes across visual scenes to find common answers.
Outcome: The proposed model can archive with 48% accuracy on the FTC, compared with state-of-the-art models.

Similar Papers

Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images (2025.naacl-srw)

Copied to clipboard

Challenge: Existing methods to measure image common sense inconsistentness are difficult to implement because of their complexity.
Approach: They propose a visual commonsense model that leverages large vision-language models to extract atomic facts from images and a compact attention-pooling classifier to fine-tune it over encoded atomic fact.
Outcome: The proposed method outperforms existing methods on the WHOOPS! and WEIRD datasets while maintaining a compact attention-pooling classifier over encoded atomic facts.
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease.
Approach: They propose to use a multiple granularity attribute-centric benchmark and training mixture to evaluate LVLMs’ fine-grained visual comprehension ability.
Outcome: The proposed model improves on LLaVa-1.5, InstructBLIP and GPT-4V and demonstrates that they struggle to generate descriptive visual attributes based on a concept that appears within an input image despite their prominent zero-shot image captioning ability.
A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) are hardly comprehensively evaluated for their cognitive abilities.
Approach: They propose to evaluate high-level cognitive abilities of Large Vision-Language Models (LVLMs) using images with rich semantics.
Outcome: The proposed evaluation benchmark consists of 251 images along with comprehensive annotations.
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in large language models have demonstrated their strong performance on IQ test questions, achieving high scores across many languages.
Approach: They propose a dataset to evaluate cognitive multimodal reasoning and problem-solving skills of large models.
Outcome: The proposed dataset contains 2,728 multiple-choice questions and 4,642 images spanning 26 categories.
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for enhancing understanding and reasoning abilities in graphbased tasks focus on specific graph types or tasks, posing challenges in designing versatile systems suitable for various tasks and graphs across diverse domains.
Approach: They propose a structure-aware fine-tuning framework to enhance LVLMs with structure learning abilities through three self-supervised learning tasks.
Outcome: Extensive evaluations on 14 LVLMs reveal that LVLs are weak in basic graph understanding and reasoning tasks, particularly those concerning relational or structurally complex information.
ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense (2023.findings-emnlp)

Copied to clipboard

Challenge: a vision-language model with commonsense knowledge can reason beyond common sense . however, pre-trained vision-linguistic models are incapable of interpreting counter-intuitive content .
Approach: They introduce a probing dataset to evaluate vision-language models' reasoning abilities . they use images that defy commonsense knowledge to test their reasoning abilities.
Outcome: The proposed dataset evaluates whether pre-trained vision-language models can reason beyond common sense . it contains images that defy commonsense knowledge with regards to color, shape, material, size and position .
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for visual commonsense reasoning (VCR) use pre-trained large language models and pre-training visionlanguage models.
Approach: They propose a collaborative approach where pre-trained LLMs serve as problem classifiers to analyze problem category and either use VLMs to answer directly or actively instruct LLM to gather relevant visual elements to support potential commonsense inferences.
Outcome: The proposed approach outperforms all other methods without in-domain fine-tuning on two VCR benchmark datasets.
A Corpus for Reasoning about Natural Language Grounded in Photographs (P19-1)

Copied to clipboard

Challenge: a dataset for visual reasoning with natural language and images is available.
Approach: They propose a dataset for joint reasoning about natural language and images . they crowdsource 107,292 examples of English sentences paired with web photographs .
Outcome: The proposed dataset combines 107,292 examples of English sentences with web photographs . Qualitative analysis shows the data requires compositional joint reasoning .
Natural Language Rationales with Full-Stack Visual Reasoning: From Pixels to Semantic Frames to Commonsense Graphs (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models that use natural language rationales provide intuitive, higher-level explanations that are easily understandable by humans.
Approach: They propose a model that generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs.
Outcome: The proposed model generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs.
Beyond Language: Learning Commonsense from Images for Reasoning (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing commonsense reasoning methods use raw texts to perform data representation and answer prediction tasks.
Approach: They propose a novel approach to learn commonsense from images instead of limited raw texts or costly knowledge bases.
Outcome: The proposed approach outperforms language-based methods on commonsense reasoning problems on two commonsence reasoning problems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations