Geo-Aware Image Caption Generation (2020.coling-main)

Copied to clipboard

Challenge: Standard image caption generation systems do not take contextual information or world knowledge into account.
Approach: They propose to build an image-specific representation of the geographic context and adapt the caption generation network to produce appropriate geographic names in the image descriptions.
Outcome: The proposed system achieves significant improvements on a dataset that contains contextualized captions and geographic metadata and improves BLEU, ROUGE, METEOR and CIDEr scores.

Similar Papers

Informative Image Captioning with External Sources of Information (P19-1)

Copied to clipboard

Challenge: Current captioning models are trained to generate captions that only contain common object names, thus falling short on an important “informativeness” dimension.
Approach: They propose a mechanism for integrating image information and fine-grained labels into a caption that describes the image in a fluent and informative manner.
Outcome: The proposed model integrates image information with fine-grained labels to produce fluent captions . it can control the appearance of these labels in the output, resulting in fluent and informative captions.
Entity-aware Image Caption Generation (D18-1)

Copied to clipboard

Challenge: Existing image captioning approaches generate generic descriptions of visual content and ignore background information.
Approach: They propose a task which generates informative image captions using images and hashtags as input.
Outcome: The proposed model outperforms unimodal baselines significantly with evaluation metrics on a dataset from Flickr.
CapOnImage: Context-driven Dense-Captioning on Image (2022.emnlp-main)

Copied to clipboard

Challenge: Existing image captioning systems generate narrative captions for images, which are spatially detached from the image in presentation.
Approach: They propose a task called captioning on image which generatesense captions at different locations of the image based on contextual information.
Outcome: The proposed model achieves the best results in both captioning accuracy and diversity aspects.
Focus! Relevant and Sufficient Context Selection for News Image Captioning (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent work only coarsely leverages the article to extract the necessary context, which makes it difficult for models to identify relevant events and named entities.
Approach: They propose to use a vision and language retrieval model CLIP to localize the visually grounded entities in the news article and then capture the non-visual entities via an open relation extraction model.
Outcome: The proposed model significantly improves on existing models and achieves state-of-the-art on multiple benchmarks.
Bridging by Word: Image Grounded Vocabulary Construction for Visual Captioning (P19-1)

Copied to clipboard

Challenge: Existing research on image captioning generates frequent n-grams with irrelevant words.
Approach: They propose to construct an image-grounded vocabulary incorporating visual information and relations among words into the decoding process directly.
Outcome: The proposed framework is compared with state-of-the-art models on MS COCO and Flickr30k and shows that it is more efficient than existing models.
Improving Image Captioning with Better Use of Caption (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to image captioning focus on visual attention, but many do not.
Approach: They propose a framework that explores semantics available in captions and leverages that to enhance both image representation and caption generation.
Outcome: The proposed framework outperforms baselines on the MSCOCO dataset and is state-of-the-art under a wide range of evaluation metrics.
Image Caption Generation for News Articles (2020.coling-main)

Copied to clipboard

Challenge: Existing work on news-image captioning requires a joint understanding of image and text.
Approach: They propose a Transformer model that integrates text and image modalities and attends to textual features from visual features in generating a caption.
Outcome: The proposed model outperforms the state-of-the-art model and improves the quality of news-image captions.
TIGEr: Text-to-Image Grounding for Image Caption Evaluation (D19-1)

Copied to clipboard

Challenge: Existing metrics based on text-level comparisons fail to assess the quality of captions produced by machines.
Approach: They propose to use a machine-learned text-image grounding model to measure the accuracy of machine-generated captions and their correlation with human judgments.
Outcome: The proposed metric has higher consistency with human judgments and is more accurate than existing metrics.
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for image captioning generate generic captions that are limited in capturing nuanced visual details.
Approach: They propose attention-guided image captioning which amplifies visual regions directly in the feature space to guide caption generation.
Outcome: The proposed approach matches or surpasses state-of-the-art models while achieving faster inference.
What Makes for Good Image Captions? (2025.findings-emnlp)

Copied to clipboard

Challenge: a formal information-theoretic framework is developed for image captioning . the pyramid of captions is a method that generates enriched captions by integrating local and global visual information.
Approach: They propose a formal information-theoretic framework for image captioning . they propose 'Pyramid of Captions' method that generates enriched captions .
Outcome: The proposed framework provides a flexible foundation for analyzing and optimizing image captioning systems across diverse task requirements.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations