Challenge: a new dataset for tweet summarization is available for free.
Approach: They propose a dataset for tweet summarization that uses human annotations to evaluate extractive summarizing methods.
Outcome: The proposed dataset includes six events collected from Twitter . human-annotated gold-standard references facilitate evaluation, the study shows .

Similar Papers

TWEETSUM: Event oriented Social Summarization Dataset (2020.coling-main)

Copied to clipboard

Challenge: Developing social summarization systems is becoming more and more critical . but, the publicly available and high-quality large scale social summaries are rare .
Approach: They propose to build a social summarization dataset using twitter's hot events . they collect user relations, hashtags and user profiles to evaluate their summarizing methods .
Outcome: The proposed dataset is based on a dataset from twitter with 12 real world hot events with 44,034 tweets and 11,240 users.
WikiSum: Coherent Summarization Dataset for Efficient Human-Evaluation (2021.acl-short)

Copied to clipboard

Challenge: Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems .
Approach: They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language.
Outcome: The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature.
TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification (2020.findings-emnlp)

Copied to clipboard

Challenge: Modern NLP systems are typically ill-equipped when applied to noisy user-generated text.
Approach: They propose a new evaluation framework consisting of seven Twitter-specific classification tasks.
Outcome: The proposed framework is based on seven heterogeneous Twitter-specific classification tasks.
SumPubMed: Summarization Dataset of PubMed Scientific Articles (2021.acl-srw)

Copied to clipboard

Challenge: Existing summarization models that can extract the top few lines of news articles fail to summarize long documents.
Approach: They constructed a scientific summarization dataset from MEDLINE articles from the PubMed archive to address this problem.
Outcome: The proposed model outperforms existing models on news article summarization datasets and shows that it is more efficient to extract the top few lines.
The State and Fate of Summarization Datasets: A Survey (2025.naacl-long)

Copied to clipboard

Challenge: Summarization is the task of shortening a text while preserving the most important information it contains.
Approach: They propose a novel ontology covering sample properties, collection methods and distribution covering sample characteristics, collection method and distribution.
Outcome: The proposed ontology covers sample properties, collection methods and distribution, and can be used to streamline future research into a more coherent body of work.
ForumSum: A Multi-Speaker Conversation Summarization Dataset (2021.findings-emnlp)

Copied to clipboard

Challenge: Abstractive summarization quality has been improved but there is a lack of data for conversation summarizing applications.
Approach: They propose to build a conversation summarization dataset with human written summaries from internet forums.
Outcome: The proposed dataset can be easily expanded to improve conversation summarization applications.
Creation and evaluation of timelines for longitudinal user posts (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods for segmenting user posts into timelines improve quality and cost of manual annotation.
Approach: They propose a set of methods for segmenting longitudinal user posts into timelines likely to contain interesting moments of change in a user’s behaviour based on their online posting activity.
Outcome: The proposed framework is able to evaluate two different social media datasets and compares with existing models.
Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects - A Survey (2024.findings-acl)

Copied to clipboard

Challenge: scholarly attention has turned to the development of text summarization methods that are more closely tailored and controlled to align with specific objectives and user needs.
Approach: They formalize a controllable text summarization task and categorize controllability attributes according to their shared characteristics and objectives.
Outcome: The proposed method is tailored to meet the specific intent and needs of users.
EntSUM: A Data Set for Entity-Centric Extractive Summarization (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for controllable summarization fail to generate entity-centric summaries.
Approach: They propose to use a human-annotated data set EntSUM to generate controllable summarization with a focus on named entities as the aspects to control.
Outcome: The proposed data set shows that existing methods fail to generate entity-centric summaries.
What Have We Achieved on Text Summarization? (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for text summarization have been investigated, but there are still gaps between them and human professionals.
Approach: They analyze 8 major sources of errors on 10 representative summarization models manually.
Outcome: Aiming to gain more understanding of summarization systems with respect to their strengths and limitations on a fine-grained syntactic and semantic level, we use 8 major sources of errors on 10 representative summarizing models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations