Challenge: WordsEye is a system for automatically converting natural language text into 3D scenes representing the meaning of that text.
Approach: They evaluate WordsEye's output vs. simple search methods to find a picture to illustrate a sentence.
Outcome: The WordsEye system produced imaginative sentences and realistic sentences . the results show that wordsEyed produced better results than standard image search engines .

Similar Papers

TIGEr: Text-to-Image Grounding for Image Caption Evaluation (D19-1)

Copied to clipboard

Challenge: Existing metrics based on text-level comparisons fail to assess the quality of captions produced by machines.
Approach: They propose to use a machine-learned text-image grounding model to measure the accuracy of machine-generated captions and their correlation with human judgments.
Outcome: The proposed metric has higher consistency with human judgments and is more accurate than existing metrics.
CLAIR: Evaluating Image Captions with Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing measures for image caption evaluation fail to capture dimensions of similarity . a novel method that leverages the zero-shot language modeling capabilities of large language models (LLMs) demonstrates a stronger correlation with human judgments of caption quality compared to existing measures.
Approach: They propose a method that leverages the zero-shot language modeling capabilities of large language models to evaluate captions.
Outcome: The proposed method shows a stronger correlation with human judgments of caption quality compared to other measures.
ImaginE: An Imagination-Based Automatic Evaluation Metric for Natural Language Generation (2023.findings-eacl)

Copied to clipboard

Challenge: Existing evaluation methods for natural language generation rely on token-level or embedding-level comparisons with text references.
Approach: They propose to use text-to-image generator to generate an image as the embodied imagination for the text snippet and compute the imagination similarity using contextual embeddings.
Outcome: The proposed metric improves existing evaluation metrics’ correlations with human similarity judgments in both reference-based and reference-free scenarios.
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? (2025.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to evaluate image captions are English-centric, despite improvements in the CLIPScore metric . however, there are no available benchmarks for multilingual captioning evaluation .
Approach: They propose to use machine-translated and machine-repurposed datasets to evaluate CLIPScore variants in multilingual settings.
Outcome: The proposed evaluation strategies are based on machine-translated and human judgements.
Transparent Human Evaluation for Image Captioning (2022.naacl-main)

Copied to clipboard

Challenge: Recent work has demonstrated that image captioning is a complex task that requires a large amount of human input.
Approach: They develop a human evaluation protocol for image captioning models based on machine- and human-generated captions on the MSCOCO dataset.
Outcome: The proposed model improves CLIPScore, a recent metric that uses image features, and improves human judgments because it is more sensitive to recall.
Assessing the State of the Art in Scene Segmentation (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in scene segmentation have made it difficult to detect scenes in literary texts.
Approach: They propose to modify existing models to improve detection of scenes in literary texts . they propose to use a training sample generation scheme to alleviate this problem .
Outcome: The proposed model is more robust to different types of texts, while its overall performance is slightly worse than that of BERT-based models.
Are Scene Graphs Good Enough to Improve Image Captioning? (2020.aacl-main)

Copied to clipboard

Challenge: Existing image captioning models rely on object detection features to generate image descriptions, but they are noisy.
Approach: They propose to use scene graphs to introduce information about object relations into captioning to improve image descriptions.
Outcome: The proposed model improves image caption quality by 3.3 CIDEr compared to a strong Bottom-Up Top-Down baseline.
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in text-to-image models have demonstrated remarkable capabilities in image synthesis.
Approach: They analyze the critical role of caption precision and recall in text-to-image model training.
Outcome: The proposed model trains with synthetic captions that show similar behavior to those trained on human-annotated captions.
Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale (2020.coling-main)

Copied to clipboard

Challenge: a few popular metrics are still used to evaluate language generation systems despite their known limitations.
Approach: They propose to use automatic metrics to evaluate language generation systems . they show that they prefer system outputs to human-authored texts .
Outcome: The proposed metrics are insensitive to correct translations of rare words and can yield high scores when given a single sentence as system output for the entire test set.
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations