Defoiling Foiled Image Captions (N18-2)

Copied to clipboard

Challenge: Existing models for vision-to-language tasks do not understand images sufficiently, despite their impressive performance.
Approach: They propose to use explicit object information to solve foiled caption detection problem . they propose to replace a word in caption with a semantically similar word .
Outcome: The proposed model achieves state-of-the-art on a recently published dataset with scores exceeding those achieved by humans on the task.

Similar Papers

Removing Word-Level Spurious Alignment between Images and Pseudo-Captions in Unsupervised Image Captioning (2021.eacl-main)

Copied to clipboard

Challenge: Unsupervised image captioning is a challenging task that requires manual annotation.
Approach: They propose a simple gating mechanism that is trained to align image features with the most reliable words in pseudo-captions.
Outcome: The proposed method outperforms the previous methods without complex learning objectives.
Informative Image Captioning with External Sources of Information (P19-1)

Copied to clipboard

Challenge: Current captioning models are trained to generate captions that only contain common object names, thus falling short on an important “informativeness” dimension.
Approach: They propose a mechanism for integrating image information and fine-grained labels into a caption that describes the image in a fluent and informative manner.
Outcome: The proposed model integrates image information with fine-grained labels to produce fluent captions . it can control the appearance of these labels in the output, resulting in fluent and informative captions.
Object Counts! Bringing Explicit Detections Back into Image Captioning (N18-1)

Copied to clipboard

Challenge: Existing approaches to image captioning use explicit object detectors as an intermediate step, but they bypass the explicit detection phase and instead generate captions directly from image embeddings.
Approach: They argue that explicit detections provide rich semantic information and can thus be used as an interpretable representation to better understand why end-to-end image captioning systems work well.
Outcome: The proposed methods can be used to understand why end-to-end captioning systems work well.
Fine-grained Image Captioning with CLIP Reward (2022.findings-naacl)

Copied to clipboard

Challenge: Modern image captioning models are usually trained with text similarity objectives . reference captions often describe only the most salient objects in images .
Approach: They propose to use CLIP to calculate multi-modal similarity and use it as a reward function . they propose a simple finetuning strategy to improve grammar that does not require extra text annotation.
Outcome: The proposed model generates more distinctive captions than the CIDEroptimized model on text-to-image retrieval and fineCapEval.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning (P18-1)

Copied to clipboard

Challenge: Practical applications of automatic image description systems include leveraging descriptions for image indexing or retrieval, and helping those with visual impairments by transforming visual signals into information that can be communicated via text-to-speech technology.
Approach: They propose to extract and filter image caption annotations from billions of webpages and use them to train models.
Outcome: The proposed model architectures perform better when trained on the Conceptual Captions dataset.
Entity-aware Image Caption Generation (D18-1)

Copied to clipboard

Challenge: Existing image captioning approaches generate generic descriptions of visual content and ignore background information.
Approach: They propose a task which generates informative image captions using images and hashtags as input.
Outcome: The proposed model outperforms unimodal baselines significantly with evaluation metrics on a dataset from Flickr.
Pragmatic Inference with a CLIP Listener for Contrastive Captioning (2023.findings-acl)

Copied to clipboard

Challenge: a new method for contrastive captioning generates discriminative captions that distinguish target images from very similar alternative distractor images.
Approach: They propose a pragmatic inference procedure that formulates captioning as a reference game between a speaker and a listener.
Outcome: The proposed method outperforms previous methods for discriminative captioning by 11% to 15% accuracy in human evaluations.
What Makes for Good Image Captions? (2025.findings-emnlp)

Copied to clipboard

Challenge: a formal information-theoretic framework is developed for image captioning . the pyramid of captions is a method that generates enriched captions by integrating local and global visual information.
Approach: They propose a formal information-theoretic framework for image captioning . they propose 'Pyramid of Captions' method that generates enriched captions .
Outcome: The proposed framework provides a flexible foundation for analyzing and optimizing image captioning systems across diverse task requirements.
Pragmatically Informative Image Captioning with Character-Level Inference (N18-2)

Copied to clipboard

Challenge: a neural image captioner and a Rational Speech Acts (RSA) model are pragmatically informative . previous attempts to combine RSA with neural image-captioning require an inference which normalizes over the entire set of possible utterances.
Approach: They propose a neural image captioner with a Rational Speech Acts model to make it pragmatically informative.
Outcome: The proposed system outperforms a non-pragmatic baseline and word-level RSA captioner on a word-based model.
Probing Image-Language Transformers for Verb Understanding (2021.findings-acl)

Copied to clipboard

Challenge: Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning.
Approach: They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech.
Outcome: The proposed model trains on a manually-annotated and smaller dataset does better on the task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations