Challenge: X-VLM models lack "fine-grained" understanding of relationships, verbs and numbers in images . pretraining on large-scale image–text data from the Web has facilitated rapid progress on many vision-and-language tasks .
Approach: They investigate models that outperform other baselines on fine-grained data . they highlight importance of novel losses and rich data sources for learning fine-grain skills .
Outcome: The proposed model outperforms baseline models on four fine-grained benchmarks . the model outpersforms other baseline models and even degrades performance .

Similar Papers

Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease.
Approach: They propose to use a multiple granularity attribute-centric benchmark and training mixture to evaluate LVLMs’ fine-grained visual comprehension ability.
Outcome: The proposed model improves on LLaVa-1.5, InstructBLIP and GPT-4V and demonstrates that they struggle to generate descriptive visual attributes based on a concept that appears within an input image despite their prominent zero-shot image captioning ability.
What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases (2024.naacl-long)

Copied to clipboard

Challenge: Vision-language models have broad competence that is difficult to evaluate . current evaluation benchmarks focus on only assessing one or a few capabilities .
Approach: They perform a large-scale transfer learning experiment to discover latent VL skills from data.
Outcome: The results suggest that factor analysis can identify reasonable yet surprising VL skill factors . the results contribute to the design of balanced and broad-coverage vision-language evaluation methods.
MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent mobile AI agents based on VLMs lack basic mobile capabilities due to their pre-trained nature.
Approach: They propose a mobile AI agent based on VLMs that includes additional pre-training stages to enhance both intra- and inter-UI understanding.
Outcome: The proposed model outperforms existing VLMs on the Chinese mobile dataset Mobile3M .
ProgressLM: Towards Progress Reasoning in Vision-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing models for task progress estimation lack long-horizon and dynamic reasoning . estimating how much of a task has been completed requires long-term reasoning based on partial information.
Approach: They propose a benchmark for evaluating progress reasoning from a single observation . they instantiate a two-stage paradigm that combines episodic retrieval with mental simulation .
Outcome: The proposed benchmark improves on 14 VLMs on a small scale and shows common failure patterns.
On the Fine-Grained Planning Abilities of VLM Web Agents (2025.findings-emnlp)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) have shown promise as web agents, yet their planning has been overlooked.
Approach: They propose to examine VLMs’ ability to understand temporal relationships within web contexts and assess plans of actions across diverse scenarios.
Outcome: The proposed models exhibit limited performance in the above skills and are not reliable to function as web agents.
Vision-Language Pretraining: Current Trends and the Future (2022.acl-tutorials)

Copied to clipboard

Challenge: Recent vision-language models are being used for downstream tasks that require large datasets and supervised datasets.
Approach: They focus on recent vision-language pretraining paradigms and their strengths and shortcomings . they compare the different family of models used for vision- language pretraining .
Outcome: This paper provides the background on image–language datasets, benchmarks, and modeling innovations before the multimodal pretraining area.
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model (2025.acl-long)

Copied to clipboard

Challenge: Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language.
Approach: They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data .
Outcome: The proposed model outperforms existing models in 14 tasks and 56 languages.
What’s “up” with vision-language models? Investigating their struggle with spatial reasoning (2023.emnlp-main)

Copied to clipboard

Challenge: Recent work has re-surfaced a concern that has long plagued vision-language models: poor performance on simple tasks like attribute attachment, counting, etc.
Approach: They evaluate 18 vision-language models and find they perform poorly on VQAv2 . they find that popular vision-linguistic pretraining corpora lack reliable data for learning spatial relationships .
Outcome: The new models are compared with existing datasets on what'sup and visual-language models . they achieve 56% accuracy on the new benchmarks compared to 99% for humans .
Does Vision-and-Language Pretraining Improve Lexical Grounding? (2021.findings-emnlp)

Copied to clipboard

Challenge: Large pretrained language models (LMs) have been criticized for lack of grounding, i.e., connecting words to their meanings in the physical world.
Approach: They compare vision-and-language (VL) models trained jointly on text and image or video data to find out how they compare to text-only counterparts.
Outcome: The proposed model outperforms the text-only variants on a commonsense question answering task.
VLURes: Benchmarking Long-Text Grounding and Cross-Lingual Robustness in Vision Language Models (2026.findings-acl)

Copied to clipboard

Challenge: ***VLURes** provides a practical testbed for long-text grounding and multilingual robustness in web-realistic agent settings.
Approach: They propose a multilingual benchmark for evaluating vision-language models under long-text grounding.
Outcome: ***VLURes** provides a testbed for long-text grounding and multilingual robustness in web-realistic agent settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations