Challenge: Existing methods for estimating uncertainty using answer likelihoods or prompt-based confidence generation often suffer from overconfidence and confirmation biases.
Approach: They propose to use Decompose and Compare Consistency (DeCC) to measure the reliability of a VLM's direct answer and indirect answers by decomposing the question into sub-questions and reasoning over the sub-answers.
Outcome: Experiments on six vision-language tasks with three VLMs show that DeCC achieves better correlation with task accuracy compared to existing methods.

Similar Papers

CAST: Cross-modal Alignment Similarity Test for Vision Language Models (2025.coling-main)

Copied to clipboard

Challenge: Vision Language Models (VLMs) are typically evaluated with Visual Question Answering tasks which assess a model’s understanding of scenes.
Approach: They propose to use visual question answering (VQA) to assess a model's understanding of scenes to probe for self-consistency across modalities.
Outcome: The proposed test does not focus on objective accuracy but rather on whether VLMs are internally consistent in their outputs.
TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Evaluating large vision-language models has focused on final-answer correctness, but this metric is often insufficient and misleading.
Approach: They propose a framework that decomposes complex multimodal tasks into Auxiliary Reasoning Sets (ARS) ARS decomposition reveals how consistently a model reasons across sub-questions with structured dependencies.
Outcome: a new framework improves diagnostic evaluation of large vision-language models . it decomposes complex multimodal tasks into auxiliary reasoning sets with structured dependencies . the framework pinpoints reasoning failures and exposes errors overlooked by standard evaluation .
MM-R3: On (In-)Consistency of Vision-Language Models (VLMs) (2025.findings-acl)

Copied to clipboard

Challenge: a flurry of research has been conducted on the performance of state-of-the-art (SoTA) Vision Language Models (VLMs) on a variety of tasks.
Approach: They propose a benchmarking tool to analyze performance of SoTA Vision Language Models (VLMs) on three tasks: Question Rephrasing, Image Restyling, and Context Reasoning.
Outcome: The proposed model achieves absolute improvements of 5.7% and 12.5% on widely used VLMs such as BLIP-2 and LLaVa 1.5M in terms of consistency over their existing counterparts.
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to quantify uncertainty are limited in vision-language models . however, current models display notable miscalibration across diverse tasks and settings .
Approach: They evaluate verbalized confidence in vision-language models using visual reasoning . they propose a prompting strategy that improves confidence alignment in multimodal settings .
Outcome: The proposed method improves confidence alignment across multimodal settings.
Confidence Improves Self-Consistency in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Modern large language models (LLMs) demonstrate strong reasoning capabilities, driven in part by their capacity to generate a sequence of intermediate reasoning steps that lead them toward a final answer.
Approach: They propose a method that performs a weighted majority vote based on confidence scores obtained directly from the model.
Outcome: The proposed method outperforms self-consistency on nine models and four datasets, reducing the required number of reasoning paths by over 40% on average.
Can VLMs Actually See and Read? A Survey on Modality Collapse in Vision-Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Vision-language models integrate textual and visual information, enabling them to process visual inputs and generate predictions.
Approach: They review work on modality collapse analysis to provide insights into the reason for this unintended behavior and review probing studies for fine-grained vision-language understanding.
Outcome: The proposed models can achieve competitive performance in vision-language tasks despite relying heavily on textual information and ignoring visual information.
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Vision-language models have demonstrated strong efficacy as visual assistants . however, evaluation of their reasoning capabilities requires a costly benchmark .
Approach: They propose a pipeline to measure the reasoning consistency of vision-language models . they propose supervised fine-tuning of VLMs and feedback from LLMs .
Outcome: The proposed framework reduces cost while ensuring the generation of a high-quality dataset.
Mixed Signals: Decoding VLMs’ Reasoning and Underlying Bias in Vision-Language Conflict (2025.findings-emnlp)

Copied to clipboard

Challenge: Vision-language models have demonstrated impressive performance by effectively integrating visual and textual information to solve complex tasks.
Approach: They build upon existing benchmarks to create five datasets containing mismatched image-text pairs and examine how they reason over visual and textual data .
Outcome: The proposed model reasoned over visual and textual data in real-world applications but not in the visual and visual descriptions.
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations (2026.acl-long)

Copied to clipboard

Challenge: Prior work has found that explanations can easily convince users that inaccurate VLM predictions are correct.
Approach: They propose to evaluate two complementary qualities of VLM-generated explanations via two quality scoring functions to improve their accuracy.
Outcome: The proposed explanations improve accuracy on the A-OKVQA, VizWiz, and MMMU-Pro tasks by 11.1%, including a 15.4% reduction in falsely believing incorrect predictions.
Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing MCQA benchmarks fail to capture the full reasoning capabilities of video language models due to selection bias.
Approach: They propose a method to reduce selection bias in video-to-text LLMs by suppressing "blind guessing" they propose 'bold' calibration technique to balance selection bias.
Outcome: The proposed method reduces selection bias and improves model performance compared to existing methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations