Semantically-Prompted Language Models Improve Visual Descriptions (2024.findings-naacl)
Copied to clipboard
| Challenge: | Language-vision models have made significant progress in zeroshot vision tasks, but lack expressive visual descriptions. |
| Approach: | They propose a new method for generating visual descriptions with pre-trained language models and semantic knowledge bases. |
| Outcome: | The proposed method improves visual descriptions and achieves strong results on image-classification datasets. |
Similar Papers
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding (2023.findings-acl)
Copied to clipboard
| Challenge: | supervised methods for vision-language tasks have been well-studied, but they lack the fine-grained information needed for semantics understanding. |
| Approach: | They propose a framework to take advantage of fine-grained information for zero-shot vision-language learning, covering multiple tasks such as VQA, SNLI-VE, and VCR. |
| Outcome: | The proposed framework outperforms previous zero-shot methods on VQA and achieves substantial improvement on SNLI-VE and VCR. |
Guiding Medical Vision-Language Models with Diverse Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations (2025.naacl-long)
Copied to clipboard
| Challenge: | Current vision-language models lack the ability to focus on specific areas designated by humans . a new framework that integrates medical entity extraction, visual prompt generation, and dataset adaptation is proposed to improve visual prompt-guided fine-tuning. |
| Approach: | They propose to use visual prompts to guide and enhance formation of region-specific attention. |
| Outcome: | The proposed framework outperforms state-of-the-art large vision-language models on medical datasets. |
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Fine-grained image classification is a challenge for vision-language models (VLMs) such as CLIP, which struggle to distinguish between semantically similar classes due to insufficient supervision for fine-grain tasks. |
| Approach: | They propose a framework that harnesses the complementary strengths of both CLIP-like and LVLMs to tackle these challenges. |
| Outcome: | The proposed framework outperforms existing models on multiple fine-grained datasets, particularly the Stanford Cars dataset. |
Visually-augmented pretrained language models for NLP tasks without images (2023.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to improve pre-trained language models lack visual commonsense and semantics. |
| Approach: | They propose a visual-augmented approach to fine-tune pre-trained language models by using retrieved or generated images instead of relying on explicit images. |
| Outcome: | The proposed approach outperforms baselines on ten tasks and consistently outperformed other approaches. |
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for rewriting text-to-image models require specialized vocabulary . a new approach uses large vision language models to optimize text-based models . |
| Approach: | They propose a prompt optimization framework that rephrases a user prompt into a text-to-image model by using large vision language models as solver and reward model. |
| Outcome: | The proposed model outperforms existing models on two popular datasets. |
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods for zero-shot video captioning focus on one key aspect of the scene and ignore the rest of the visual input. |
| Approach: | They propose a novel textual prompting strategy for zero-shot video captioning that uses a category-aware retrieval mechanism to promote prompt diversity while ensuring visual relevance. |
| Outcome: | The proposed method outperforms existing methods on in-domain and cross-domain settings. |
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Vision-Language Models have demonstrated impressive performance on vision-language reasoning tasks, but their potential for zero-shot fine-grained image classification remains underexplored. |
| Approach: | They propose a method that transforms zero-shot fine-grained image classification into a visual question-answering framework. |
| Outcome: | The proposed method outperforms the current state-of-the-art approach and outperformed existing methods. |
Enhancing Large Vision-Language Models with Ultra-Detailed Image Caption Generation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing pipelines for generating high-quality, ultra-detailed image captions are limited by the scarcity of image caption data. |
| Approach: | They propose a pipeline for generating high-quality, ultra-detailed image captions that integrates both pre-processing and post-processor stages. |
| Outcome: | The proposed pipeline improves LVLMs' perception and cognitive abilities across multiple vision-language benchmarks. |
CLIPText: A New Paradigm for Zero-shot Text Classification (2023.findings-acl)
Copied to clipboard
| Challenge: | Experimental results show that CLIP can be applied to zero-shot text classification tasks. |
| Approach: | They propose a CLIP model for zero-shot text classification that integrates prompt into CLIPText to better derive knowledge from CLIP. |
| Outcome: | The proposed model can be applied to a text-image matching problem and show that it can be used for language tasks. |
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease. |
| Approach: | They propose to use a multiple granularity attribute-centric benchmark and training mixture to evaluate LVLMs’ fine-grained visual comprehension ability. |
| Outcome: | The proposed model improves on LLaVa-1.5, InstructBLIP and GPT-4V and demonstrates that they struggle to generate descriptive visual attributes based on a concept that appears within an input image despite their prominent zero-shot image captioning ability. |