Challenge: Recent studies have proposed reference-free evaluations of image captions . however, these approaches are restrictive and favor captions with similar vocabulary but different meanings.
Approach: They propose to use reference-free metrics to evaluate image captions . they propose to combine lexical overlap and semantics to identify fine-grained errors .
Outcome: The proposed metrics struggle to identify fine-grained errors, the authors show . CLIPScore, UMIC, and PAC-S are sensitive to variations in image-relevant objects mentioned in the caption .

Similar Papers

CLIPScore: A Reference-free Evaluation Metric for Image Captioning (2021.emnlp-main)

Copied to clipboard

Challenge: Image captioning relies on reference-based automatic evaluations, but references are expensive to collect and comparing against multiple human-authored captions is insufficient.
Approach: They propose a reference-free metric that can be used for automatic caption evaluation without references.
Outcome: The proposed model outperforms existing metrics on image-text compatibility and a reference-augmented version achieves even higher correlation with human judgements.
Do Image–Text Metrics Respect Semantic Invariances? (2026.findings-acl)

Copied to clipboard

Challenge: Reference-free image–to–text evaluators are now standard for scoring image–caption alignment, yet it is unclear whether they respect semantic invariances.
Approach: They propose an invariance probe on five popular evaluators under semantics-preserving perturbations along three axes: spatial edits, object changes, and socio-linguistic framing.
Outcome: The proposed invariance probe shows that spatial edits and simple phrasing changes shift scores by ()6% on average and cause ranking flips in up to (),37% of cases.
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: N-gram-based evaluation metrics are unreliable due to low correlation to human judgments.
Approach: They propose a metric that rewards correct details and penalizes incorrect ones.
Outcome: The proposed metric matches the performance of open-source LLM-based metrics in correlation to human judgments while being far more efficient.
PR-MCS: Perturbation Robust Metric for MultiLingual Image Captioning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing image captioning metrics are vulnerable to lexical perturbations, but they are not robust to such perturbations.
Approach: They propose a perturbation-robust multilingual CLIPScore which is a reference-free image captioning metric for multiple languages.
Outcome: The proposed metric outperforms baseline metrics in capturing lexical noise of all various perturbation types in all five languages while maintaining a strong correlation with human judgments.
InfoMetIC: An Informative Metric for Reference-free Image Caption Evaluation (2023.acl-long)

Copied to clipboard

Challenge: Existing image captioning metrics provide a single score to measure caption qualities, which are less explainable and informative.
Approach: They propose an Informative Metric for Reference-free Image Caption evaluation to support this feedback . they propose to provide a text precision score, a vision recall score and an overall quality score .
Outcome: The proposed method improves on existing metrics on multiple benchmarks and compares coarse-grained scores with human judgements.
Learning-based Composite Metrics for Improved Caption Evaluation (P18-3)

Copied to clipboard

Challenge: Existing image captioning metrics focus on linguistic aspects and do not match human judgements at sentence-level.
Approach: They propose to incorporate lexical and semantic metrics as features to capture adequacy and fluency of captions at different linguistic levels.
Outcome: The proposed framework captures adequacy and fluency of captions at different linguistic levels.
UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning (2021.acl-short)

Copied to clipboard

Challenge: BERTScore and other text generation metrics do not use reference captions to evaluate image captions.
Approach: They propose a new metric which does not require reference captions to evaluate image captions . they train UMIC to discriminate negative captions via contrastive learning .
Outcome: The proposed metric has higher correlation than previous metrics that require multiple references.
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? (2025.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to evaluate image captions are English-centric, despite improvements in the CLIPScore metric . however, there are no available benchmarks for multilingual captioning evaluation .
Approach: They propose to use machine-translated and machine-repurposed datasets to evaluate CLIPScore variants in multilingual settings.
Outcome: The proposed evaluation strategies are based on machine-translated and human judgements.
Transparent Human Evaluation for Image Captioning (2022.naacl-main)

Copied to clipboard

Challenge: Recent work has demonstrated that image captioning is a complex task that requires a large amount of human input.
Approach: They develop a human evaluation protocol for image captioning models based on machine- and human-generated captions on the MSCOCO dataset.
Outcome: The proposed model improves CLIPScore, a recent metric that uses image features, and improves human judgments because it is more sensitive to recall.
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model (2024.acl-long)

Copied to clipboard

Challenge: Existing image captioning evaluation metrics do not provide an explanation for the assigned numerical score.
Approach: They propose an explainable reference-free metric to provide an explanation for captions . they introduce score smoothing to align as closely as possible with human judgment .
Outcome: The proposed metric achieves high correlations with human judgment across image captioning evaluation benchmarks and is publicly available at https://github.com/Yebin46/FLEUR.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations