Challenge: a recent study examines the evaluation of hotel highlights in the context of hotel data.
Approach: They examine evaluation of faithfulness to input data in the context of hotel highlights . they compare traditional metrics, trainable methods, and LLM-as-a-judge approaches .
Outcome: The results show that simple metrics outperform human judgments on LLM-generated summaries . the results also highlight challenges in crowdsourced evaluations.

Similar Papers

ACUEval: Fine-grained Hallucination Evaluation and Correction for Abstractive Summarization (2024.findings-acl)

Copied to clipboard

Challenge: Recent-proposed evaluation metrics for large language models have a preference-bias . however, such metrics often lack interpretability and only offer a single score .
Approach: They propose a metric that leverages the power of large language models to perform two sub-tasks: decomposing summaries into atomic content units and validating them against the source document.
Outcome: The proposed metric improves faithfulness scores on three summarization evaluation benchmarks by 3% compared to the next-best metric.
Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation (2023.acl-long)

Copied to clipboard

Challenge: Existing studies for summarization evaluation exhibit low inter-annotator agreement or lack scale.
Approach: They propose a modified summarization salience protocol based on fine-grained semantic units and a robust summarizing evaluation benchmark.
Outcome: The proposed protocol is based on fine-grained semantic units and allows for high inter-annotator agreement.
SummEval: Re-evaluating Summarization Evaluation (2021.tacl-1)

Copied to clipboard

Challenge: a lack of comprehensive studies on evaluation metrics for text summarization hinders progress . a new study aims to improve evaluation metrics that correlate with human judgments .
Approach: They propose to re-evaluate automatic evaluation metrics and share a toolkit for evaluation . they hope to promote a more complete evaluation protocol for text summarization .
Outcome: The proposed evaluation metrics are inconsistent with existing evaluation protocols.
FineSurE: Fine-grained Summarization Evaluation using LLMs (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for text summarization evaluation do not correlate well with human judgments . evaluators that use Likert scale scores are limited in their ability to perform deeper analysis.
Approach: They propose a fine-grained evaluator specifically tailored for the summarization task using large language models.
Outcome: The proposed method improves on open-source and proprietary LLMs and shows better completeness and conciseness than existing methods.
LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization (2023.eacl-main)

Copied to clipboard

Challenge: Human evaluation is labor-intensive, expensive to scale, and difficult to design.
Approach: They propose a set of guidelines for human evaluation of faithfulness in long-form summaries that address the following challenges: (1) How can we achieve high inter-annotator agreement on faithfulness scores? (2) How can our annotator minimize workload while maintaining accurate faithfulness?
Outcome: The proposed framework reduces inter-annotator variance in faithfulness scores while minimizing annotator workload while maintaining accuracy.
How Far are We from Robust Long Abstractive Summarization? (2022.emnlp-main)

Copied to clipboard

Challenge: Abstractive summarization has made tremendous progress in recent years . however, even under a short document setting, abstractive models often generate summaries that are repetitive, ungrammatical, and factually inconsistent with the source.
Approach: They perform fine-grained human annotations to evaluate long document abstractive summarization systems and develop factual consistency metrics.
Outcome: The proposed model can generate more relevant summaries but not factual ones.
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)

Copied to clipboard

Challenge: Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models .
Approach: They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets.
Outcome: The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets.
Studying Summarization Evaluation Metrics in the Appropriate Scoring Range (P19-1)

Copied to clipboard

Challenge: Existing evaluation metrics are compared based on their ability to correlate with humans, but they disagree in the higher-scoring range in which current systems operate.
Approach: They show that evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
Outcome: The evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges (2024.emnlp-main)

Copied to clipboard

Challenge: introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance.
Approach: They propose a taxonomy for organizing existing LLM-based evaluation metrics and a structured framework to understand and compare them.
Outcome: The proposed taxonomy offers a framework to understand and compare LLM-based evaluation methods.
Summarization Evaluation in the Absence of Human Model Summaries Using the Compositionality of Word Embeddings (C18-1)

Copied to clipboard

Challenge: Existing summary evaluation methods rely on multiple model summaries to evaluate quality of summary outputs.
Approach: They propose a new summary evaluation approach that does not require human model summaries . they exploit compositional capabilities of word embeddings to develop features .
Outcome: The proposed metric replicates human-generated summarization scores on data from TAC 2008 and 2009 . the features are then used to train a learning model for predicting the summary content quality in the absence of gold models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations