Challenge: Language-vision models have made significant progress in zeroshot vision tasks, but lack expressive visual descriptions.
Approach: They propose a new method for generating visual descriptions with pre-trained language models and semantic knowledge bases.
Outcome: The proposed method improves visual descriptions and achieves strong results on image-classification datasets.

Similar Papers

UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding (2023.findings-acl)

Copied to clipboard

Challenge: supervised methods for vision-language tasks have been well-studied, but they lack the fine-grained information needed for semantics understanding.
Approach: They propose a framework to take advantage of fine-grained information for zero-shot vision-language learning, covering multiple tasks such as VQA, SNLI-VE, and VCR.
Outcome: The proposed framework outperforms previous zero-shot methods on VQA and achieves substantial improvement on SNLI-VE and VCR.
Guiding Medical Vision-Language Models with Diverse Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations (2025.naacl-long)

Copied to clipboard

Challenge: Current vision-language models lack the ability to focus on specific areas designated by humans . a new framework that integrates medical entity extraction, visual prompt generation, and dataset adaptation is proposed to improve visual prompt-guided fine-tuning.
Approach: They propose to use visual prompts to guide and enhance formation of region-specific attention.
Outcome: The proposed framework outperforms state-of-the-art large vision-language models on medical datasets.
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Fine-grained image classification is a challenge for vision-language models (VLMs) such as CLIP, which struggle to distinguish between semantically similar classes due to insufficient supervision for fine-grain tasks.
Approach: They propose a framework that harnesses the complementary strengths of both CLIP-like and LVLMs to tackle these challenges.
Outcome: The proposed framework outperforms existing models on multiple fine-grained datasets, particularly the Stanford Cars dataset.
Visually-augmented pretrained language models for NLP tasks without images (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to improve pre-trained language models lack visual commonsense and semantics.
Approach: They propose a visual-augmented approach to fine-tune pre-trained language models by using retrieved or generated images instead of relying on explicit images.
Outcome: The proposed approach outperforms baselines on ten tasks and consistently outperformed other approaches.
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for rewriting text-to-image models require specialized vocabulary . a new approach uses large vision language models to optimize text-based models .
Approach: They propose a prompt optimization framework that rephrases a user prompt into a text-to-image model by using large vision language models as solver and reward model.
Outcome: The proposed model outperforms existing models on two popular datasets.
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for zero-shot video captioning focus on one key aspect of the scene and ignore the rest of the visual input.
Approach: They propose a novel textual prompting strategy for zero-shot video captioning that uses a category-aware retrieval mechanism to promote prompt diversity while ensuring visual relevance.
Outcome: The proposed method outperforms existing methods on in-domain and cross-domain settings.
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Vision-Language Models have demonstrated impressive performance on vision-language reasoning tasks, but their potential for zero-shot fine-grained image classification remains underexplored.
Approach: They propose a method that transforms zero-shot fine-grained image classification into a visual question-answering framework.
Outcome: The proposed method outperforms the current state-of-the-art approach and outperformed existing methods.
Enhancing Large Vision-Language Models with Ultra-Detailed Image Caption Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing pipelines for generating high-quality, ultra-detailed image captions are limited by the scarcity of image caption data.
Approach: They propose a pipeline for generating high-quality, ultra-detailed image captions that integrates both pre-processing and post-processor stages.
Outcome: The proposed pipeline improves LVLMs' perception and cognitive abilities across multiple vision-language benchmarks.
CLIPText: A New Paradigm for Zero-shot Text Classification (2023.findings-acl)

Copied to clipboard

Challenge: Experimental results show that CLIP can be applied to zero-shot text classification tasks.
Approach: They propose a CLIP model for zero-shot text classification that integrates prompt into CLIPText to better derive knowledge from CLIP.
Outcome: The proposed model can be applied to a text-image matching problem and show that it can be used for language tasks.
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease.
Approach: They propose to use a multiple granularity attribute-centric benchmark and training mixture to evaluate LVLMs’ fine-grained visual comprehension ability.
Outcome: The proposed model improves on LLaVa-1.5, InstructBLIP and GPT-4V and demonstrates that they struggle to generate descriptive visual attributes based on a concept that appears within an input image despite their prominent zero-shot image captioning ability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations