Language Agnostic Automatic Summarization Evaluation (2020.lrec-1)

Copied to clipboard

Challenge: Existing evaluation methods for summarization of documents have been primarily focused on the English language.
Approach: They propose to use ROUGE and PYRAMID to evaluate non-English data using English and non- English data sets.
Outcome: The proposed evaluation methods can be adapted to non-English data, and the results show that they can perform well on non- English data.

Similar Papers

Re-Evaluating Evaluation for Multilingual Summarization (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that automated evaluation approaches correlate with human ratings in English, but this is unclear for other languages.
Approach: They construct a small-scale pilot dataset containing article-summary pairs and human ratings in English, Chinese and Indonesian to measure the strength of summaries.
Outcome: The results show that standard metrics are unreliable measures of quality in Chinese and Indonesian.
Does Summary Evaluation Survive Translation to Other Languages? (2022.naacl-main)

Copied to clipboard

Challenge: a quality summarization dataset requires the production and evaluation of summaries by trained humans and machines.
Approach: They translate a summarization dataset in English and compare its performance to seven languages . they explore equivalence testing as an appropriate statistical paradigm for evaluating correlations between human and automated scoring of summaries .
Outcome: The proposed method could be used in seven languages and compares performance across measures.
A Survey on Cross-Lingual Summarization (2022.tacl-1)

Copied to clipboard

Challenge: Cross-lingual summarization is a task of generating a summary in one language for a given document in a different language.
Approach: They present a systematic review of the literature on cross-lingual summarization . they summarize previous efforts and compare them with each other .
Outcome: The proposed approach is compared with previous approaches and summarizes them to provide a deeper analysis.
Evaluating the Efficacy of Summarization Evaluation across Languages (2021.findings-acl)

Copied to clipboard

Challenge: Using multilingual summarization evaluation methods is more reliable and interpretable than manual methods.
Approach: They propose to use multilingual BERT within BERTScore to evaluate summarization evaluation metrics . they use English datasets that are not representative of modern summarizing systems .
Outcome: The proposed methods perform well across all languages, at a level above that for English.
The State and Fate of Summarization Datasets: A Survey (2025.naacl-long)

Copied to clipboard

Challenge: Summarization is the task of shortening a text while preserving the most important information it contains.
Approach: They propose a novel ontology covering sample properties, collection methods and distribution covering sample characteristics, collection method and distribution.
Outcome: The proposed ontology covers sample properties, collection methods and distribution, and can be used to streamline future research into a more coherent body of work.
Automatic Pyramid Evaluation Exploiting EDU-based Extractive Reference Summaries (D18-1)

Copied to clipboard

Challenge: Existing methods for evaluating content are not accurate because they only confirm if the summary contains small textual fragments.
Approach: They propose to transform human-made reference summaries into extractive reference sums and weight them using elementary discourse units.
Outcome: The proposed method strongly correlates with manual evaluations on DUC and TAC data sets.
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization (2025.acl-long)

Copied to clipboard

Challenge: n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear.
Approach: They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics.
Outcome: The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand.
SummEval: Re-evaluating Summarization Evaluation (2021.tacl-1)

Copied to clipboard

Challenge: a lack of comprehensive studies on evaluation metrics for text summarization hinders progress . a new study aims to improve evaluation metrics that correlate with human judgments .
Approach: They propose to re-evaluate automatic evaluation metrics and share a toolkit for evaluation . they hope to promote a more complete evaluation protocol for text summarization .
Outcome: The proposed evaluation metrics are inconsistent with existing evaluation protocols.
HOLMS: Alternative Summary Evaluation with Large Language Models (2020.coling-main)

Copied to clipboard

Challenge: Efficient document summarization requires evaluation measures that can rank a set of systems based on an average score and highlight which individual summary is better than another.
Approach: They propose a hybrid evaluation measure for document summarization called HOLMS that combines both language models pre-trained on large corpora and lexical similarity measures.
Outcome: The proposed measure outperforms ROUGE and BLEU on several extractive summarization datasets for both linguistic quality and pyramid scores.
Revisiting Automatic Evaluation of Extractive Summarization Task: Can We Do Better than ROUGE? (2022.findings-acl)

Copied to clipboard

Challenge: Existing methods to evaluate text summarization tasks using ROUGE have been criticized for lack of semantic understanding.
Approach: They propose a semantic-aware metric for extractive summarization task that is semantic-based . they use CNN/DailyMail dataset to study the new metric .
Outcome: The proposed metric is semantic-aware and shows higher correlation with human judgement and yields a large number of disagreements with the original ROUGE metric.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations