Papers with Vision
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain (2024.findings-acl)
Copied to clipboard
Liang Chen, Yichi Zhang, Shuhuai Ren, Haozhe Zhao, Zefan Cai, Yuchi Wang, Peiyi Wang, Xiangdi Meng, Tianyu Liu, Baobao Chang
| Challenge: | a new multimodal decision-making benchmark evaluates the integrated capabilities of multimodal large language models. |
| Approach: | They propose a multimodal decision-making benchmark for evaluating MLLMs . they propose an automatic evaluation protocol to assess 10 prevalent ML models . |
| Outcome: | The proposed benchmark improves performance of multimodal large language models in three scenarios . the model is required to integrate multiple capabilities to make accurate decisions . |
i-Code V2: An Autoregressive Generation Framework over Vision, Language, and Speech Data (2024.findings-naacl)
Copied to clipboard
Ziyi Yang, Mahmoud Khademi, Yichong Xu, Reid Pryzant, Yuwei Fang, Chenguang Zhu, Dongdong Chen, Yao Qian, Xuemei Gao, Yi-Ling Chen, Robert Gmyr, Naoyuki Kanda, Noel Codella, Bin Xiao, Yu Shi, Lu Yuan, Takuya Yoshioka, Michael Zeng, Xuedong Huang
| Challenge: | i-Code V2 is one of the first models capable of generating natural language from any combination of Vision, Language, and Speech data. |
| Approach: | They propose to create a model that can generate natural language from any combination of Vision, Language, and Speech data. |
| Outcome: | i-Code V2 matches or outperforms state-of-the-art single- and dual-modality baselines on 7 multimodal tasks. |
Curriculum Masking in Vision-Language Pretraining to Maximize Cross Modal Interaction (2024.naacl-long)
Copied to clipboard
| Challenge: | masked language modeling is widely used as a pretraining component in Vision and language (V+L) but performance on benchmarks has not received the attention it deserves. |
| Approach: | They propose a curriculum masking scheme that uses a parallel mask selection agent to mask tokens at a frequency proportional to the level of cross modal interaction necessary to reconstruct them. |
| Outcome: | The proposed method improves relational understanding on a wide range of V+L tasks. |
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks (2023.acl-long)
Copied to clipboard
| Challenge: | Vision and language models exploit unrobust indicators in individual modalities instead of focusing on relevant information in each modality. |
| Approach: | They propose a performance-agnostic multimodality score based on Shapley values that quantifies in which proportions a multimodal model uses individual modalities. |
| Outcome: | The proposed model can quantify in which proportions a multimodal model uses individual modalities for different tasks and datasets. |
Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing hallucination benchmarks rarely test this failure mode outside Western contexts and English. |
| Approach: | They propose a multimodal benchmark built from images spanning 17 MENA countries . they use a CFHR-based test to measure hallucination beyond raw accuracy . |
| Outcome: | The proposed model is based on images from 17 MENA countries . it measures counterfactual acceptance conditioned on correctly answering the true statement. |
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding (2026.findings-eacl)
Copied to clipboard
| Challenge: | figurative language is essential for expressing intent, emotion, and perspective . figural language is often dependent on Styles Reasoning, causing incongruities between expressions . |
| Approach: | They propose a framework that induces reasoning capabilities to compact vision–language models . figurative language is essential in expressing intent, emotion, and perspective . |
| Outcome: | The proposed framework can interpret multimodal figurative language, provide transparent reasoning traces, and generalize across multiple figurativ styles. |
Follow the Beaten Path: The Role of Route Patterns on Vision-Language Navigation Agents Generalization Abilities (2025.naacl-long)
Copied to clipboard
| Challenge: | Vision and language navigation (VLN) is a challenging task towards the creation of embodied agents. |
| Approach: | They propose a solution that combines visual and linguistic features to enable VLN . they propose augmentation of the training data to fill the gap in missing patterns . |
| Outcome: | The proposed solution fills the gap in missing patterns of training data. |
Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas (2022.acl-long)
Copied to clipboard
| Challenge: | Recent work on visual-grounded navigation has focused on indoor scenarios with sharp drops in performance when testing on unseen data. |
| Approach: | They focus on visual agent navigation in outdoor scenarios with panorama images . they find that most gain in outdoor VLN on unseen data is due to specific features . |
| Outcome: | The results show a bias to specifics of graph representations of urban environments, demanding that VLN tasks grow in scale and diversity of geographical environments. |
ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents (2026.acl-long)
Copied to clipboard
| Challenge: | Visually rich documents (VRDs) combine text, tables, and figures within complex, semantically structured layouts. |
| Approach: | They propose a multi-turn reinforcement learning framework that fine-tunes VLMs as interactive agents capable of actively navigating long, visually rich documents. |
| Outcome: | The proposed framework achieves state-of-the-art on five long-document benchmarks. |
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing studies have focused on assessing the model’s overall accuracy without evaluating it on different reasoning cases. |
| Approach: | They propose a novel idea to identify and improve multi-modal multi-hop reasoning in VQA by using two new language prompts to find a reasoning path to reach its answer. |
| Outcome: | The proposed model improves multi-modal multi-hop reasoning in visual question answering (VQA) it finds that the proposed model is easy to answer, simply demanding “single-hop” reasoning, whereas only a few questions require “multi-hop.” |
Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene Graphs (2023.emnlp-main)
Copied to clipboard
Roei Herzig, Alon Mendelson, Leonid Karlinsky, Assaf Arbelle, Rogerio Feris, Trevor Darrell, Amir Globerson
| Challenge: | Vision and language models (VLMs) have demonstrated remarkable zero-shot (ZS) performance in a variety of tasks. |
| Approach: | They propose to integrate structured annotations into visual and textual representations to improve VLMs' understanding of compositional scenes. |
| Outcome: | The proposed method improves VLMs on multiple VL datasets with only a mild degradation in ZS capabilities. |
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision–Language Modeling (2026.findings-acl)
Copied to clipboard
| Challenge: | FTibSuite provides an end-to-end training-and-evaluation workflow for vision–language models . Tibetan is underserved due to the lack of infrastructure for reproducible training and evaluation. |
| Approach: | They propose a resource-centric workflow for Tibetan VLMs that provides an end-to-end training-and-evaluation workflow and human-verified multimodal annotations. |
| Outcome: | FTibSuite provides an end-to-end training-and-evaluation workflow and human-verified multimodal annotations. |
Lost in Embeddings: Information Loss in Vision–Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Experiments reveal connectors substantially distort the local geometry of visual representations, with k-nearest neighbors diverging by 40–60% post-projection, correlating with degradation in retrieval performance. |
| Approach: | They propose two approaches to examine and quantify information loss by analyzing latent representation space. |
| Outcome: | The proposed model improves retrieval performance by analyzing changes in k-nearest neighbor relationships between image representations before and after projection. |