Papers with Summarization
Proceedings of the 2nd Workshop on New Frontiers in Summarization (D19-54)
Copied to clipboard
| Challenge: | EMNLP 2017 is a workshop on enhancing natural language processing's ability to produce concise, fluent summaries. |
| Approach: | the workshop provides a forum for cross-fertilization of ideas towards automatic summarization . four invited speakers will be present at the workshop . |
| Outcome: | the workshop aims to provide a forum for cross-fertilization of ideas towards automatic summarization. |
Contrastive Data and Learning for Natural Language Processing (2022.naacl-tutorials)
Copied to clipboard
| Challenge: | Current NLP models heavily rely on effective representation learning algorithms. |
| Approach: | This tutorial introduces contrastive learning and provides an introduction to the techniques. |
| Outcome: | This tutorial provides an introduction to the fundamentals of contrastive learning approaches and the theory behind them. |
Fine-grained Factual Consistency Assessment for Abstractive Summarization Models (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have shown that around 30% of the summaries generated by abstractive summarization models contain factual errors. |
| Approach: | They propose a fine-grained two-stage Fact Consistency assessment framework for summarization models that uses fine-grain consistency reasoning to find subtle clues to identify whether a model-generated summary is consistent with the original document. |
| Outcome: | The proposed framework improves on the state-of-the-art models and distinguishes detailed differences better. |
WARP-Text: a Web-Based Tool for Annotating Relationships between Pairs of Texts (C18-2)
Copied to clipboard
| Challenge: | Existing tools for annotating pairs of texts do not support detailed pairwise annotation. |
| Approach: | They present an open-source web-based tool for annotating relationships between pairs of texts . they propose to use WARP-Text to create multi-layer annotations and custom definitions . |
| Outcome: | The proposed tool can be used by project managers and annotators. |
HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder Models (2022.emnlp-main)
Copied to clipboard
| Challenge: | Abstractive summarization systems implicitly encode “decisions” about summary properties, but these are not enforced. |
| Approach: | They propose a new summarization architecture that extends existing models to a mixture-of-experts version with multiple decoders. |
| Outcome: | The proposed architecture outperforms baseline models in obtaining stylistically-diverse summaries by sampling from individual decoders or their mixtures. |
STRASS: A Light and Effective Method for Extractive Summarization Based on Sentence Embeddings (P19-2)
Copied to clipboard
| Challenge: | Summarization is a costly and timedemanding task. |
| Approach: | They propose an extractive text summarization method which leverages the semantic information in existing sentence embedding spaces. |
| Outcome: | The proposed method performs similarly to state-of-the-art extractive methods with effective training and inference time. |
RLHF Algorithms Ranked: An Extensive Evaluation Across Diverse Tasks, Rewards, and Hyperparameters (2025.emnlp-industry)
Copied to clipboard
Lucas Spangher, Rama Kumar Pasumarthi, Nick Masiewicki, William F. Arnold, Aditi Kaushal, Dale Johnson, Peter Grabowski, Eugene Ie
| Challenge: | Proximal Policy Optimization (PPO) has fallen out of favor for Large Language Models (LLMs), but its complexity and inefficiency have spurred the investigation of simpler alternatives. |
| Approach: | They evaluate 17 RLHF algorithms on two benchmarks, OpenAI’s TL;DR Summarization and Anthropic’s Helpfulness / Harmlessness. |
| Outcome: | The proposed methods are based on OpenAI’s TL;DR Summarization and Anthropic’s Helpfulness / Harmlessness benchmarks with two different reward models and a Rules based reward model. |
ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text Mining (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent large-scale language models show remarkable achievements in key NLP tasks such as Question Answering and Text Summarization. |
| Approach: | They propose a domain-specific pre-trained Vietnamese language model that outperforms the general domain language models. |
| Outcome: | The proposed model outperforms the general domain language models in Vietnamese datasets while outperforming the general-domain language models. |
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs (2025.naacl-short)
Copied to clipboard
Forrest Sheng Bao, Miaoran Li, Renyi Qu, Ge Luo, Erana Wan, Yujia Tang, Weisi Fan, Manveer Singh Tamber, Suleman Kazi, Vivek Sourabh, Mike Qi, Ruixuan Tu, Chenyu Xu, Matthew Gonzales, Ofer Mendelevitch, Amin Ahmad
| Challenge: | Existing evaluations of hallucinations in large language models suffer from a lack of diversity and recency in the LLM and LLM families considered. |
| Approach: | They propose a summarization hallucination benchmark that challenges models to disagree on hallucines . they use models to generate answers or summaries from textual input . |
| Outcome: | The proposed model combines the best of 10 modern LLMs with ground truth annotations. |
Neural Label Search for Zero-Shot Multi-Lingual Extractive Summarization (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods to translate sentences to other languages using heuristics are challenging. |
| Approach: | They propose a model that learns hierarchical weights for different sets of labels and applies them to other languages to translate them. |
| Outcome: | The proposed model can translate English datasets to other languages and obtain different sets of labels again using heuristics. |
A Cascade Approach to Neural Abstractive Summarization with Content Selection and Fusion (2020.aacl-main)
Copied to clipboard
| Challenge: | Existing systems that perform content selection and surface realization are not able to provide sufficient training data for news summarization. |
| Approach: | They propose to use a cascade architecture to perform content selection and surface realization together to generate abstracts. |
| Outcome: | The proposed architecture outperforms or outranks existing systems in terms of content selection and surface realization. |
Summarizing Medical Conversations via Identifying Important Utterances (2020.coling-main)
Copied to clipboard
| Challenge: | Applying natural language processing (NLP) techniques to the medical field is a prevailing trend nowadays and has great potential in many applications, such as key information extraction in medical literature. |
| Approach: | They propose to use a hierarchical encoder-tagger model to generate medical conversation summarization by identifying important utterances. |
| Outcome: | The proposed model outperforms baseline models and models and adds conversation-related features to improve performance. |
SQuALITY: Building a Long-Document Summarization Dataset the Hard Way (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing summarization datasets often have issues that seriously limit their usability. |
| Approach: | They propose a faster but more straightforward approach to developing summarization benchmark data . they use a protocol that hires highly-qualified contractors to read stories and write original summaries from scratch . |
| Outcome: | The proposed protocol is faster but more straightforward than scraping summaries from everyday text. |
Improving Factuality of Abstractive Summarization without Sacrificing Summary Quality (2023.acl-short)
Copied to clipboard
| Challenge: | Recent studies have shown that most abstractive summarization models are unfaithful and suffer from a wide range of hallucination. |
| Approach: | They propose a candidate summary generation and ranking technique to improve summary factuality without sacrificing quality. |
| Outcome: | The proposed method shows that the model trained using the proposed method improves on factuality and similarity-based metrics without conflicting with the model. |
FELIX: Flexible Text Editing Through Tagging and Insertion (2020.findings-emnlp)
Copied to clipboard
| Challenge: | FELIX is efficient in low-resource settings and fast at inference time, while being capable of modeling flexible input-output transformations. |
| Approach: | They propose a flexible text-editing approach that decomposes a text-generating task into two sub-tasks: tagging and insertion. |
| Outcome: | The proposed model is efficient in low-resource settings and fast at inference time while being capable of modeling flexible input-output transformations. |
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation are expensive and time-consuming. |
| Approach: | They propose a framework that utilizes LLMs to generate synthetic evaluation datasets . they propose meta-correlation to measure alignment between metric rankings and human benchmarks based on synthetic data . |
| Outcome: | The proposed framework achieves meta-correlations exceeding 0.9 in multilingual QA and replaces human judgment with synthetic evaluation datasets. |
An Investigation of Evaluation Methods in Automatic Medical Note Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent studies show that doctors can save significant amounts of time when using automatic note generation. |
| Approach: | They propose task-specific metrics for automatic note generation from medical conversation summarization and generation, including knowledge-graph embedding-based metrics, customized model-based measures with domain-specific weights, and ensemble metrics. |
| Outcome: | The proposed evaluation metrics are compared to existing models and can have different behaviors on different types of clinical notes datasets. |
MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answering (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for answering time-sensitive questions lack temporal reasoning . existing methods struggle with these time-intensive questions, authors say . |
| Approach: | They propose a temporal-based question-answering framework that integrates temporal perturbations and gold evidence labels into a question processing framework. |
| Outcome: | The proposed framework outperforms baseline retrieval methods in retrieval performance. |
ALIGNMEET: A Comprehensive Tool for Meeting Annotation, Alignment, and Evaluation (2022.lrec-1)
Copied to clipboard
| Challenge: | Summarization is a challenging problem, and it is difficult to create, correct, and evaluate the summaries manually. |
| Approach: | They propose an open-source tool for meeting annotation, alignment, and evaluation . the tool aims to provide an efficient and clear interface for fast annotation . |
| Outcome: | The proposed tool is open-source and installable from PyPI. |
Can you Summarize my learnings? Towards Perspective-based Educational Dialogue Summarization (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Increasing use of virtual tutors has allowed for more efficient, personalized, and interactive AI-based learning experiences. |
| Approach: | They propose a task of Multi-modal Perspective based Dialogue Summarization (MM-PerSumm) that summarizes educational dialogues from three unique perspectives: the Student, the Tutor, and a Generic viewpoint. |
| Outcome: | The proposed model can summarize educational dialogues from three perspectives, while student-oriented summaries should distill learning points, track progress, and suggest scope for improvement. |
SummaCoz: A Dataset for Improving the Interpretability of Factual Consistency Detection for Summarization (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Summarization is an important application of Large Language Models. |
| Approach: | They integrate human-annotated and model-generated natural language explanations to elucidate how a summary deviates and becomes inconsistent with its source article. |
| Outcome: | The proposed model provides rationales for its judgments and improves its accuracy significantly. |
ETPC - A Paraphrase Identification Corpus Annotated with Extended Paraphrase Typology and Negation (L18-1)
Copied to clipboard
| Challenge: | Extended Paraphrase Typology addresses limitations of existing typologies . extended typology provides better means for evaluation and error analysis . |
| Approach: | a new typology copes with non-paraphrase pairs in paraphrase identification corpora, a paper proposes . a large corpus annotated with atomic paraphrase types is the largest to date . |
| Outcome: | The Extended Paraphrase Typology (EPT) and the Extended Typology Paraphrase Corpus (ETPC) address practical limitations of existing paraphrase typologies. |
Klexikon: A German Dataset for Joint Summarization and Simplification (2022.lrec-1)
Copied to clipboard
| Challenge: | Traditionally, Text Simplification is a monolingual translation task where individual sentences are "translated" into a simplified version. |
| Approach: | They propose to use a dataset to jointly simplify long source documents by combining sentences from a source and their simplified counterparts. |
| Outcome: | The proposed system can summarize and simplify long source documents using almost 2,900 documents. |
HighRES: Highlight-based Reference-less Evaluation of Summarization (P19-1)
Copied to clipboard
| Challenge: | Existing methods for summarizing documents are inconsistent due to the difficulty of manual evaluation. |
| Approach: | They propose a method where summaries are evaluated by multiple annotators against the source document via manually highlighted salient content. |
| Outcome: | The proposed method improves inter-annotator agreement while highlighting differences among systems. |
CoCoA: Confidence- and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing contrastive decoding methods that handle conflict lack adaptability and can degrade performance in low conflict settings. |
| Approach: | They propose a token-level algorithm for principled conflict resolution and enhanced faithfulness that resolves conflict by utilizing confidence-aware measures and the generalized divergence between parametric and contextual distributions. |
| Outcome: | The proposed algorithm achieves 9.2 points on average in QA, summarization, and long-form question answering (LFQA) benchmarks and improves factuality by 2.5 points on the key benchmarks. |
Bias in News Summarization: Measures, Pitfalls and Corpora (2024.findings-acl)
Copied to clipboard
| Challenge: | Pretrained large language models can reproduce harmful social biases in constrained settings, such as summarization. |
| Approach: | They propose a method to generate input documents with carefully controlled demographic attributes and then apply it to a controlled setting. |
| Outcome: | The proposed method allows to generate input documents with carefully controlled demographic attributes while working with real-world input documents. |
CCSum: A Large-Scale and High-Quality Dataset for Abstractive News Summarization (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing datasets for supervised news summarization contain considerable amount of noise and expensive training data. |
| Approach: | They propose a large-scale and high-quality dataset for supervised abstractive news summarization containing 1.3 million training samples. |
| Outcome: | The proposed dataset is more factual and informative than established summarization datasets. |
DHP Benchmark: Are LLMs Good NLG Evaluators? (2025.findings-naacl)
Copied to clipboard
Yicheng Wang, Jiayi Yuan, Yu-Neng Chuang, Zhuoer Wang, Yingchi Liu, Mark Cusick, Param Kulkarni, Zhengping Ji, Yasser Ibrahim, Xia Hu
| Challenge: | Large Language Models (LLMs) are increasingly serving as evaluators in Natural Language Generation (NLG) tasks. |
| Approach: | They propose a framework that measures the discernment of Large Language Models (LLMs) across diverse NLG tasks. |
| Outcome: | The proposed framework provides quantitative discernment scores for LLMs across four NLG tasks. |
A Repository of Corpora for Summarization (L18-1)
Copied to clipboard
| Challenge: | Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task. |
| Approach: | They propose a repository containing corpora available to train and evaluate automatic summarization systems. |
| Outcome: | The proposed system is based on a repository of corpora available for summarization tasks. |
How to Compare Things Properly? A Study of Argument Relevance in Comparative Question Answering (2025.acl-long)
Copied to clipboard
Irina Nikishina, Saba Anwar, Nikolay Dolgov, Maria Manina, Daria Ignatenko, Artem Shelmanov, Chris Biemann
| Challenge: | Comparative Question Answering (CQA) is a task that involves processing information and diverse viewpoints. |
| Approach: | They construct a dataset of arguments annotated with their relevance and use it to answer comparative questions. |
| Outcome: | The proposed dataset contains arguments annotated with their relevance and enables precise traceability and faithfulness. |
Are Large Language Models In-Context Personalized Summarizers? Get an iCOPERNICUS Test Done! (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have succeeded in summarizing information in contexts but saliency is subject to user preferences. |
| Approach: | They propose a framework that measures saliency using user reading histories and contrast in user profiles. |
| Outcome: | The proposed framework evaluates state-of-the-art LLMs on their ICL performance and shows that they lack true ICPL. |
From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to align large language models with human preferences suffer from inconsistent scoring and suboptimal alignment. |
| Approach: | They propose a dual-consistency framework that aligns partial sequences with human preferences. |
| Outcome: | The proposed framework significantly reduces granularity discrepancies and improves GPT-4 evaluation scores. |
CSTree-SRI: Introspection-Driven Cognitive Semantic Tree for Multi-Turn Question Answering over Extra-Long Contexts (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have achieved remarkable success in natural language processing (NLP), particularly in single-turn question answering (QA) on short-text. |
| Approach: | They propose a framework that captures logical correlations across chunks of ELC and maintains coherence of multi-turn Questions. |
| Outcome: | The proposed framework is able to capture logical correlations across chunks of ELC and maintain coherence of multi-turn Questions. |
Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for multimodal summarization often inject shallow visual features into deep models, leading to representational mismatches and weak cross-modal grounding. |
| Approach: | They propose a framework that performs text summarization and representative image selection . a deep visual processor aligns the visual encoder with the language model at corresponding depths . |
| Outcome: | The proposed framework produces more accurate, visually grounded summaries and selects more representative images. |