Training for Diversity in Image Paragraph Captioning (D18-1)

Copied to clipboard

Challenge: Existing image captioning models have a lack of diversity between sentences . current models have limited their effectiveness due to repetitive paragraphs .
Approach: They propose to apply sequence-level training to image paragraph captioning models . they find that standard self-critical training produces poor results .
Outcome: The proposed training improves on the Visual Genome dataset with no architectural changes.

Similar Papers

Image Embedding Sampling Method for Diverse Captioning (2025.emnlp-main)

Copied to clipboard

Challenge: Currently, large-scale captioning models are less accessible for resource-constrained applications such as mobile devices and assistive technologies.
Approach: They propose a training-free framework that enhances caption diversity and informativeness by explicitly attending to distinct image regions using a comparably small VLM as the backbone.
Outcome: The proposed framework achieves comparable performance to larger models on MSCOCO, Flickr30k, and Nocaps test datasets while maintaining strong image-caption relevancy and semantic integrity with the human-annotated captions.
Cross-Modal Similarity-Based Curriculum Learning for Image Captioning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing image captioning approaches treat image-caption pairs indistinctly without considering the differences in their learning difficulties.
Approach: They propose a pretrained vision–language model that measures cross-modal similarity and a model that uses cross-module similarity to measure the difficulty of captioning.
Outcome: The proposed model achieves superior performance and competitive convergence speed to baselines without incurring additional training costs.
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in text-to-image models have demonstrated remarkable capabilities in image synthesis.
Approach: They analyze the critical role of caption precision and recall in text-to-image model training.
Outcome: The proposed model trains with synthetic captions that show similar behavior to those trained on human-annotated captions.
Improving Image Captioning with Better Use of Caption (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to image captioning focus on visual attention, but many do not.
Approach: They propose a framework that explores semantics available in captions and leverages that to enhance both image representation and caption generation.
Outcome: The proposed framework outperforms baselines on the MSCOCO dataset and is state-of-the-art under a wide range of evaluation metrics.
CapOnImage: Context-driven Dense-Captioning on Image (2022.emnlp-main)

Copied to clipboard

Challenge: Existing image captioning systems generate narrative captions for images, which are spatially detached from the image in presentation.
Approach: They propose a task called captioning on image which generatesense captions at different locations of the image based on contextual information.
Outcome: The proposed model achieves the best results in both captioning accuracy and diversity aspects.
Improving Visual Question Answering by Referring to Generated Paragraph Captions (P19-1)

Copied to clipboard

Challenge: Empirical results show that paragraph captions help answer more visual questions .
Approach: They propose a visual and textual question answering model which uses paragraph captions as input . they use cross-attention to extract related information, then consensus to fuse the inputs .
Outcome: Empirical results show that paragraph captions help answer more visual questions . the proposed model significantly improves the baseline model .
Learning from Children: Improving Image-Caption Pretraining via Curriculum (2023.findings-acl)

Copied to clipboard

Challenge: Image-caption pretraining is a difficult problem as it requires multiple concepts (nouns) from captions to be aligned to multiple objects in images.
Approach: They propose a curriculum learning framework that uses images to align multiple concepts to multiple objects in an image.
Outcome: The proposed learning framework improves over pretraining from scratch, using a pretrained image or/and text encoder, low data regime etc.
Neural Caption Generation for News Images (L18-1)

Copied to clipboard

Challenge: Existing methods for automatic caption generation of images are lacking in the field of image-related applications.
Approach: They propose a method for automatically generating captions for news images . they propose several deep neural network architectures built upon Recurrent Neural Networks .
Outcome: The proposed method outperforms a traditional method on a BBC News dataset using automatic evaluation and human evaluation.
Uncovering Limitations in Text-to-Image Generation: A Contrastive Approach with Structured Semantic Alignment (2023.findings-emnlp)

Copied to clipboard

Challenge: a new method for text-to-image generation models is proposed to address these limitations . SSA focuses on learning structured semantic embeddings across different modalities .
Approach: They propose a method to evaluate text-to-image generation models using structured semantic embeddings . they propose to learn mutated prompts by substituting words with equivalent or nonequivalent alternatives .
Outcome: The proposed method improves the measurement of semantic consistency of text-to-image generation models.
Pretrained Image-Text Models are Secretly Video Captioners (2025.naacl-short)

Copied to clipboard

Challenge: Current video captioning methods often incorporate intricate designs tailored to video inputs.
Approach: They adapt an image-based captioning model to address dynamic video sequences without modifications.
Outcome: The proposed model outperforms specialised captioning systems on major benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations