Papers with FGVC
Semantically-Prompted Language Models Improve Visual Descriptions (2024.findings-naacl)
Copied to clipboard
| Challenge: | Language-vision models have made significant progress in zeroshot vision tasks, but lack expressive visual descriptions. |
| Approach: | They propose a new method for generating visual descriptions with pre-trained language models and semantic knowledge bases. |
| Outcome: | The proposed method improves visual descriptions and achieves strong results on image-classification datasets. |
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease. |
| Approach: | They propose to use a multiple granularity attribute-centric benchmark and training mixture to evaluate LVLMs’ fine-grained visual comprehension ability. |
| Outcome: | The proposed model improves on LLaVa-1.5, InstructBLIP and GPT-4V and demonstrates that they struggle to generate descriptive visual attributes based on a concept that appears within an input image despite their prominent zero-shot image captioning ability. |