Pragmatically Informative Image Captioning with Character-Level Inference (N18-2)

Copied to clipboard

Challenge: a neural image captioner and a Rational Speech Acts (RSA) model are pragmatically informative . previous attempts to combine RSA with neural image-captioning require an inference which normalizes over the entire set of possible utterances.
Approach: They propose a neural image captioner with a Rational Speech Acts model to make it pragmatically informative.
Outcome: The proposed system outperforms a non-pragmatic baseline and word-level RSA captioner on a word-based model.

Similar Papers

Pragmatic Inference with a CLIP Listener for Contrastive Captioning (2023.findings-acl)

Copied to clipboard

Challenge: a new method for contrastive captioning generates discriminative captions that distinguish target images from very similar alternative distractor images.
Approach: They propose a pragmatic inference procedure that formulates captioning as a reference game between a speaker and a listener.
Outcome: The proposed method outperforms previous methods for discriminative captioning by 11% to 15% accuracy in human evaluations.
Pragmatic Issue-Sensitive Image Captioning (2020.findings-emnlp)

Copied to clipboard

Challenge: Issue-Sensitive Image Captioning (ISIC) is a new approach to image captioning . high-quality captions are shaped by the communicative goal of identifying the target image .
Approach: They propose to use image partitions to control image caption generation to produce descriptive captions.
Outcome: The proposed model can be extended to include image partitions and image partitioning.
Informative Image Captioning with External Sources of Information (P19-1)

Copied to clipboard

Challenge: Current captioning models are trained to generate captions that only contain common object names, thus falling short on an important “informativeness” dimension.
Approach: They propose a mechanism for integrating image information and fine-grained labels into a caption that describes the image in a fluent and informative manner.
Outcome: The proposed model integrates image information with fine-grained labels to produce fluent captions . it can control the appearance of these labels in the output, resulting in fluent and informative captions.
What Makes for Good Image Captions? (2025.findings-emnlp)

Copied to clipboard

Challenge: a formal information-theoretic framework is developed for image captioning . the pyramid of captions is a method that generates enriched captions by integrating local and global visual information.
Approach: They propose a formal information-theoretic framework for image captioning . they propose 'Pyramid of Captions' method that generates enriched captions .
Outcome: The proposed framework provides a flexible foundation for analyzing and optimizing image captioning systems across diverse task requirements.
CLAIR: Evaluating Image Captions with Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing measures for image caption evaluation fail to capture dimensions of similarity . a novel method that leverages the zero-shot language modeling capabilities of large language models (LLMs) demonstrates a stronger correlation with human judgments of caption quality compared to existing measures.
Approach: They propose a method that leverages the zero-shot language modeling capabilities of large language models to evaluate captions.
Outcome: The proposed method shows a stronger correlation with human judgments of caption quality compared to other measures.
Entity-aware Image Caption Generation (D18-1)

Copied to clipboard

Challenge: Existing image captioning approaches generate generic descriptions of visual content and ignore background information.
Approach: They propose a task which generates informative image captions using images and hashtags as input.
Outcome: The proposed model outperforms unimodal baselines significantly with evaluation metrics on a dataset from Flickr.
On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization (2022.findings-emnlp)

Copied to clipboard

Challenge: Combining visual modality with pretrained language models has been effective for descriptive tasks such as image captioning.
Approach: They ask: do multimodal models combine visual and visual adapted language models? they find that CLIP image representations and scaling of language models do not consistently improve self-rationalization in multimodal tasks.
Outcome: The proposed model types do not consistently improve self-rationalization in multimodal tasks.
Bridging by Word: Image Grounded Vocabulary Construction for Visual Captioning (P19-1)

Copied to clipboard

Challenge: Existing research on image captioning generates frequent n-grams with irrelevant words.
Approach: They propose to construct an image-grounded vocabulary incorporating visual information and relations among words into the decoding process directly.
Outcome: The proposed framework is compared with state-of-the-art models on MS COCO and Flickr30k and shows that it is more efficient than existing models.
IC3: Image Captioning by Committee Consensus (2023.emnlp-main)

Copied to clipboard

Challenge: Traditionally, image captioning models are trained to generate a single “best’ (most like a reference) image caption.
Approach: They propose a method to generate a single caption that captures high-level details from several annotator viewpoints.
Outcome: The proposed method outperforms baseline SOTA models and improves the performance of automated recall systems by up to 84%.
Evaluating pragmatic abilities of image captioners on A3DS (2023.acl-short)

Copied to clipboard

Challenge: Evaluating grounded neural language models with respect to pragmatic qualities such as truthfulness, contrastivity and overinformativity remains a challenge in absence of data collected from humans.
Approach: They propose to use an open source image-text dataset to evaluate pragmatic abilities of grounded neural language models with respect to pragmatic qualities such as truthfulness, contrastivity and over-informativity.
Outcome: The proposed model develops human-like pragmatic abilities with respect to truthfulness, contrastivity and over-informativity for specific features.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations