Challenge: a growing number of documents are needed for multi-document summarization.
Approach: They propose crowdsourcing to evaluate intrinsic and extrinsic quality of extractive text summaries . they conduct intensive comparative crowdsourcing and laboratory experiments .
Outcome: The proposed crowdsourcing task evaluates intrinsic and extrinsic quality of extractive text summaries.

Similar Papers

Investigating Crowdsourcing Protocols for Evaluating the Factual Consistency of Summaries (2022.naacl-main)

Copied to clipboard

Challenge: Existing pre-trained summarization models produce text that is factually inconsistent with the input.
Approach: They present a scale-based scale for Likert rating and a scoring algorithm for Best-Worst Scaling to improve crowdsourcing reliability.
Outcome: The proposed model is more reliable than existing models on two news summarization datasets.
Is Summary Useful or Not? An Extrinsic Human Evaluation of Text Summaries on Downstream Tasks (2024.lrec-main)

Copied to clipboard

Challenge: a recent study focused on intrinsic evaluation, which assesses the quality of summaries, e.g. coherence, fluency, and informativeness, but it focused on task-based extrinsic evaluation to determine the usefulness of summarizations.
Approach: They incorporate three downstream tasks to measure the usefulness of summaries . they find that fine-tuned models produce more useful summary across all three tasks .
Outcome: The proposed model produces more useful summaries across all three tasks compared to zero-shot models . human evaluation provides more reliable performance assessment compared with automatic methods .
SummEval: Re-evaluating Summarization Evaluation (2021.tacl-1)

Copied to clipboard

Challenge: a lack of comprehensive studies on evaluation metrics for text summarization hinders progress . a new study aims to improve evaluation metrics that correlate with human judgments .
Approach: They propose to re-evaluate automatic evaluation metrics and share a toolkit for evaluation . they hope to promote a more complete evaluation protocol for text summarization .
Outcome: The proposed evaluation metrics are inconsistent with existing evaluation protocols.
Intrinsic Evaluation of Summarization Datasets (2020.emnlp-main)

Copied to clipboard

Challenge: Almost all popular summarization datasets do not come with inherent quality assurance guarantees.
Approach: They propose to use 5 metrics to evaluate quality of summarization datasets . they find that data usage in recent summarizing research is inconsistent with the properties of the data.
Outcome: The proposed metrics can be inexpensive heuristics for detecting generically low quality examples.
Facet-Aware Evaluation for Extractive Summarization (2020.acl-main)

Copied to clipboard

Challenge: lexical overlap is a common evaluation metric for extractive summarization, but recent studies reveal its limitations.
Approach: They propose a facet-aware evaluation setup for better assessment of information coverage in extractive summaries.
Outcome: The proposed evaluation setup improves human correlation with extractive summarization datasets and improves comparative analysis.
Re-evaluating Evaluation in Text Summarization (2020.emnlp-main)

Copied to clipboard

Challenge: Automated evaluation metrics are an essential part of the development of text-generation tasks such as summarization.
Approach: They propose to use top-scoring system outputs to assess the reliability of automatic evaluation metrics for text summarization.
Outcome: The proposed evaluation method is based on human judgments from 25 top-scoring neural summarization systems.
Crowdsourcing Lightweight Pyramids for Manual Summary Evaluation (N19-1)

Copied to clipboard

Challenge: Manual evaluation methods are perceived as insufficient due to the high cost of the Pyramid method and the required expertise.
Approach: They propose a crowdsourced method that compares system summaries to references and uses crowdsourced scripts to analyze the results.
Outcome: The proposed method shows higher correlation relative to the original Pyramid method.
What Have We Achieved on Text Summarization? (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for text summarization have been investigated, but there are still gaps between them and human professionals.
Approach: They analyze 8 major sources of errors on 10 representative summarization models manually.
Outcome: Aiming to gain more understanding of summarization systems with respect to their strengths and limitations on a fine-grained syntactic and semantic level, we use 8 major sources of errors on 10 representative summarizing models.
Exploring the Cost-Effectiveness of Perspective Taking in Crowdsourcing Subjective Assessment: A Case Study of Toxicity Detection (2025.naacl-long)

Copied to clipboard

Challenge: toxicity evaluation tasks require annotations to accurately reflect opinions of subgroups . toxicity tasks require annotators to take the opinions of a subgroup simultaneously .
Approach: They propose to use perspective taking to obtain opinions from subgroups . they propose to prompt annotators to take perspectives of contrasting subgroup simultaneously .
Outcome: The proposed approach can be cost-effective and improve quality under limited budget.
Effective Crowdsourcing for a New Type of Summarization Task (N18-2)

Copied to clipboard

Challenge: Currently, summarization research focuses on summarizing the entire text, but in practice, readers are often interested in only one aspect of the document or conversation.
Approach: They propose a new task where the goal is to summarize a particular aspect of a document.
Outcome: The proposed task is based on a crowdsourced data collection workflow that allows users to collect high-quality summaries.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations