Pragmatically Informative Image Captioning with Character-Level Inference (N18-2)
Copied to clipboard
| Challenge: | a neural image captioner and a Rational Speech Acts (RSA) model are pragmatically informative . previous attempts to combine RSA with neural image-captioning require an inference which normalizes over the entire set of possible utterances. |
| Approach: | They propose a neural image captioner with a Rational Speech Acts model to make it pragmatically informative. |
| Outcome: | The proposed system outperforms a non-pragmatic baseline and word-level RSA captioner on a word-based model. |
Similar Papers
Pragmatic Inference with a CLIP Listener for Contrastive Captioning (2023.findings-acl)
Copied to clipboard
| Challenge: | a new method for contrastive captioning generates discriminative captions that distinguish target images from very similar alternative distractor images. |
| Approach: | They propose a pragmatic inference procedure that formulates captioning as a reference game between a speaker and a listener. |
| Outcome: | The proposed method outperforms previous methods for discriminative captioning by 11% to 15% accuracy in human evaluations. |
Pragmatic Issue-Sensitive Image Captioning (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Issue-Sensitive Image Captioning (ISIC) is a new approach to image captioning . high-quality captions are shaped by the communicative goal of identifying the target image . |
| Approach: | They propose to use image partitions to control image caption generation to produce descriptive captions. |
| Outcome: | The proposed model can be extended to include image partitions and image partitioning. |
Informative Image Captioning with External Sources of Information (P19-1)
Copied to clipboard
| Challenge: | Current captioning models are trained to generate captions that only contain common object names, thus falling short on an important “informativeness” dimension. |
| Approach: | They propose a mechanism for integrating image information and fine-grained labels into a caption that describes the image in a fluent and informative manner. |
| Outcome: | The proposed model integrates image information with fine-grained labels to produce fluent captions . it can control the appearance of these labels in the output, resulting in fluent and informative captions. |
What Makes for Good Image Captions? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a formal information-theoretic framework is developed for image captioning . the pyramid of captions is a method that generates enriched captions by integrating local and global visual information. |
| Approach: | They propose a formal information-theoretic framework for image captioning . they propose 'Pyramid of Captions' method that generates enriched captions . |
| Outcome: | The proposed framework provides a flexible foundation for analyzing and optimizing image captioning systems across diverse task requirements. |
CLAIR: Evaluating Image Captions with Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing measures for image caption evaluation fail to capture dimensions of similarity . a novel method that leverages the zero-shot language modeling capabilities of large language models (LLMs) demonstrates a stronger correlation with human judgments of caption quality compared to existing measures. |
| Approach: | They propose a method that leverages the zero-shot language modeling capabilities of large language models to evaluate captions. |
| Outcome: | The proposed method shows a stronger correlation with human judgments of caption quality compared to other measures. |
Entity-aware Image Caption Generation (D18-1)
Copied to clipboard
| Challenge: | Existing image captioning approaches generate generic descriptions of visual content and ignore background information. |
| Approach: | They propose a task which generates informative image captions using images and hashtags as input. |
| Outcome: | The proposed model outperforms unimodal baselines significantly with evaluation metrics on a dataset from Flickr. |
On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Combining visual modality with pretrained language models has been effective for descriptive tasks such as image captioning. |
| Approach: | They ask: do multimodal models combine visual and visual adapted language models? they find that CLIP image representations and scaling of language models do not consistently improve self-rationalization in multimodal tasks. |
| Outcome: | The proposed model types do not consistently improve self-rationalization in multimodal tasks. |
Bridging by Word: Image Grounded Vocabulary Construction for Visual Captioning (P19-1)
Copied to clipboard
| Challenge: | Existing research on image captioning generates frequent n-grams with irrelevant words. |
| Approach: | They propose to construct an image-grounded vocabulary incorporating visual information and relations among words into the decoding process directly. |
| Outcome: | The proposed framework is compared with state-of-the-art models on MS COCO and Flickr30k and shows that it is more efficient than existing models. |
IC3: Image Captioning by Committee Consensus (2023.emnlp-main)
Copied to clipboard
| Challenge: | Traditionally, image captioning models are trained to generate a single “best’ (most like a reference) image caption. |
| Approach: | They propose a method to generate a single caption that captures high-level details from several annotator viewpoints. |
| Outcome: | The proposed method outperforms baseline SOTA models and improves the performance of automated recall systems by up to 84%. |
Evaluating pragmatic abilities of image captioners on A3DS (2023.acl-short)
Copied to clipboard
| Challenge: | Evaluating grounded neural language models with respect to pragmatic qualities such as truthfulness, contrastivity and overinformativity remains a challenge in absence of data collected from humans. |
| Approach: | They propose to use an open source image-text dataset to evaluate pragmatic abilities of grounded neural language models with respect to pragmatic qualities such as truthfulness, contrastivity and over-informativity. |
| Outcome: | The proposed model develops human-like pragmatic abilities with respect to truthfulness, contrastivity and over-informativity for specific features. |