Papers with summarization

300 papers
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias (2024.naacl-short)

Copied to clipboard

Challenge: Position bias is a tendency of a model unfairly prioritizing information from certain parts of the input text over others, leading to undesirable behavior.
Approach: They propose to measure position bias in large language models for zero-shot summarization tasks by measuring position bias.
Outcome: The proposed model performance and position biases lead to new insights and discussion on zero-shot summarization tasks.
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand (2026.acl-long)

Copied to clipboard

Challenge: Existing summarization systems struggle to address diverse linguistic and cognitive barriers among general readers.
Approach: They propose a multi-agent framework that integrates template-based planning with an iterative feedback loop guided by simulated readers and domain expert revision to address comprehension barriers such as unknown terms, missing contexts, and confusing sentences.
Outcome: The proposed framework improves readability and factuality across multiple datasets and human evaluations show that it is more accessible to a wide range of readers.
Can LLMs Learn Macroeconomic Narratives from Social Media? (2025.findings-naacl)

Copied to clipboard

Challenge: Existing evaluation strategies for analyzing economic data with narratives are limited due to the complexity of the interplay of numerous factors and the difficulty in isolating causal relationships.
Approach: They propose to use two Twitter datasets to capture economy-related narratives and use them to construct models using large language models.
Outcome: The proposed models are able to predict macroeconomic fluctuations using the extracted or extracted narratives in two Twitter datasets.
jTLEX: a Java Library for TimeLine EXtraction (2023.eacl-demo)

Copied to clipboard

Challenge: Timeline EXtraction library provides Java implementation of TimeML annotations and tools for programmatic manipulation of Timeline graphs.
Approach: jTLEX provides a Java implementation of TimeLine EXtraction algorithm and utilities for programmatic manipulation of TimeML graphs.
Outcome: jTLEX provides a Java implementation of the TimeLine EXtraction algorithm, along with utilities for programmatic manipulation of TimeML graphs.
Catching Attention with Automatic Pull Quote Selection (2020.coling-main)

Copied to clipboard

Challenge: Using a novel task, we advocate automatic pull quote selection to engage readers with thought-provoking articles . pull quotes increase enjoyment and readability, shape reader perceptions, and facilitate learning.
Approach: They propose a task that automatically selects pull quotes from articles with more salient presentation.
Outcome: The proposed task differs from similar tasks such as summarization and clickbait identification by several aspects.
An End-to-End Dialogue Summarization System for Sales Calls (2022.naacl-industry)

Copied to clipboard

Challenge: Summarizing sales calls is a routine task performed manually by salespeople.
Approach: They propose a production system which combines generative models fine-tuned for customer-agent setting, with a human-in-the-loop user experience for an interactive summary curation process.
Outcome: The proposed system can handle training data scarcity and privacy constraints in an industrial setting.
A Call for Clarity in Beam Search: How It Works and When It Stops (2024.lrec-main)

Copied to clipboard

Challenge: Empirical results show that a modified beam decoding implementation improves decoding performance of strong, neural language generation models.
Approach: They propose a modification to a beam decoding implementation that generalizes the stopping criterion and provides flexibility to the depth of search.
Outcome: The proposed method improves decoding performance of strong models on news text summarization and machine translation over diverse language pairs with negligible inference slowdown.
SETSum: Summarization and Visualization of Student Evaluations of Teaching (2022.naacl-demo)

Copied to clipboard

Challenge: Student Evaluations of Teaching (SETs) are used in colleges and universities to assess student perceptions about their courses.
Approach: They propose a system that leverages sentiment analysis, aspect extraction, summarization and visualization techniques to provide organized illustrations of SET findings to instructors and other reviewers.
Outcome: The proposed system can be used by 10 professors from diverse departments to analyze SET results.
Attention Temperature Matters in Abstractive Summarization Distillation (2022.acl-long)

Copied to clipboard

Challenge: Recent progress of abstractive text summarization relies on large pre-trained sequence-to-sequence Transformer models, which are computationally expensive.
Approach: They propose to distill large Transformer summarization models into smaller ones with minimal performance loss by manipulating attention temperatures in Transformers.
Outcome: The proposed method outperforms vanilla pseudo-labeling based methods on three summarization datasets and is shorter and more abstractive.
A Thorough Evaluation of Task-Specific Pretraining for Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Previous work has used task-agnostic pretraining methods like masked language models or corrupted span prediction to improve performance on downstream tasks.
Approach: They propose to use a task-agnostic pretraining to improve on low-resource tasks.
Outcome: The proposed model can predict extracted gap sentences on summarization with a low resource and zero shot setup.
WikiAsp: A Dataset for Multi-domain Aspect-based Summarization (2021.tacl-1)

Copied to clipboard

Challenge: Existing aspects-based summarization models are domain-specific due to large differences in the type of aspects for different domains.
Approach: They propose a large-scale dataset for multi-domain aspect-based summarization using Wikipedia articles from 20 different domains.
Outcome: The proposed model is based on Wikipedia articles from 20 different domains and uses the section titles and boundaries of each article as a proxy for aspect annotation.
NarraSum: A Large-Scale Dataset for Abstractive Narrative Summarization (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on summarizing news documents or structured documents.
Approach: They propose to use a large-scale narrative summarization dataset to encourage research . they find there is a performance gap between humans and the models on NarraSum .
Outcome: The proposed dataset shows that humans and state-of-the-art models perform poorly when summarizing a narrative . it contains 122K narratives collected from synopses of movies and TV episodes with diverse genres .
Constructing a Dataset for Hallucination Detection in Japanese Summarization with Fine-grained Faithfulness Labels (2026.eacl-srw)

Copied to clipboard

Challenge: Large language models (LLMs) can generate fluent text, but the quality of generated content depends on its consistency with the given input.
Approach: They constructed a Japanese evaluation dataset for hallucination detection in summarization by manually annotating sentence-level faithfulness labels in LLM-generated summaries of Japanese documents.
Outcome: The proposed model can detect hallucinations in Japanese documents by annotating faithfulness labels in Japanese summaries.
Detecting and Mitigating Challenges in Zero-Shot Video Summarization with Video LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Video Large Language Models (VLLMs) exhibit impressive zero-shot capabilities in video analysis, but their performance varies significantly depending on the LLM prompt, the characteristics of the video, and the properties of the training data and LLM architecture.
Approach: They propose to use Chain-of-Thought prompting to inject knowledge extracted by external, lightweight models into video summarization benchmarks to evaluate their performance.
Outcome: The proposed solutions improve summarization performance by injecting knowledge extracted by external, lightweight models.
CoDesc: A Large Code–Description Parallel Dataset (2021.findings-acl)

Copied to clipboard

Challenge: Existing models for natural language and programming languages are lagging behind due to a lack of large datasets and benchmarks.
Approach: They present a large parallel dataset of Java methods and natural language descriptions that is used to train deep neural models.
Outcome: The proposed dataset improves code summarization and code search by 22% and opens up possibilities for pretrained language models for Java.
A Multi-Level Optimization Framework for End-to-End Text Augmentation (2022.tacl-1)

Copied to clipboard

Challenge: Existing methods for text augmentation perform data augmentation and downstream tasks separately.
Approach: They propose a framework to perform text augmentation and the downstream task end-to-end.
Outcome: The proposed framework performs text augmentation and the downstream task end-to-end on a text classification dataset.
AREDSUM: Adaptive Redundancy-Aware Iterative Sentence Ranking for Extractive Document Summarization (2021.eacl-main)

Copied to clipboard

Challenge: Existing studies on redundancy are focused on salience alone.
Approach: They propose to combine salience and novelty to score redundancy in extractive summarization systems . they also propose to balance saliance and redundancies by scoring redundants first .
Outcome: Empirical results show that AREDSUM-CTX scores salience first, then learns to balance saliency and redundancy.
Enhancing Large Language Models for Scientific Multimodal Summarization with Multimodal Output (2025.coling-industry)

Copied to clipboard

Challenge: Scientific publications are becoming more multimedia, containing both text and visual content.
Approach: They propose a framework for Scientific Multimodal Summarization with Multimodal Output . it leverages the power of large language models and extends its capability to cross-modal understanding .
Outcome: The proposed framework outperforms uni- and multi-modality methods on two new datasets . it leverages the power of large language models and extends its capability to cross-modal understanding .
Summary Explorer: Visualizing the State of the Art in Text Summarization (2021.emnlp-demo)

Copied to clipboard

Challenge: Automatic text summarization is the task of generating a summary of a long text by condensing it to its most important parts.
Approach: They propose a tool to visually explore document summarization systems based on three well-known summary quality criteria .
Outcome: The proposed tool compiles outputs of 55 state-of-the-art document summarization approaches and visually explores them during a qualitative assessment.
RELexED: Retrieval-Enhanced Legal Summarization with Exemplar Diversity (2025.findings-naacl)

Copied to clipboard

Challenge: Current approaches to legal summarization struggle with content theme deviation and inconsistent writing styles due to the content of the source document.
Approach: They propose a retrieval-augmented framework that utilizes exemplar summaries along with the source document to guide the model.
Outcome: The proposed model outperforms models that do not utilize exemplars and those that rely on similarity-based exemplar selection.
Texar: A Modularized, Versatile, and Extensible Toolkit for Text Generation (P19-3)

Copied to clipboard

Challenge: Texar is an open-source text generation toolkit that supports a broad set of text generation tasks.
Approach: They introduce Texar, an open-source text generation toolkit that supports text generation tasks.
Outcome: Texar supports machine translation, summarization, dialog, content manipulation, and more.
On-device System of Compositional Multi-tasking in Large Language Models (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing approaches to generative AI for large language models struggle when executing complex tasks simultaneously.
Approach: They propose a novel approach tailored specifically for compositional multi-tasking scenarios . they add a learnable projection layer on top of the combined summarization and translation adapters.
Outcome: The proposed approach performs well and is fast in both cloud-based and on-device implementations.
Decontextualization: Making Sentences Stand-Alone (2021.tacl-1)

Copied to clipboard

Challenge: Taking excerpts of text can be problematic, as key pieces may not be explicit in a local window.
Approach: They define a problem of sentence decontextualization by rewriting a sentence to be interpretable out of context while preserving its meaning.
Outcome: The proposed method can be used in question answering and document understanding tasks.
WikiSum: Coherent Summarization Dataset for Efficient Human-Evaluation (2021.acl-short)

Copied to clipboard

Challenge: Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems .
Approach: They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language.
Outcome: The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature.
CASPER: Bridging Discrete and Continuous Prompt Optimization through Feedback-Guided Gradient Descent (2026.eacl-industry)

Copied to clipboard

Challenge: Existing pipelines for generative tasks require extensive manual effort and domain expertise to achieve task-optimal performance.
Approach: They propose a framework bridging discrete and continuous prompt optimization through feedback-guided gradient descent in embedding space.
Outcome: The proposed framework bridges discrete and continuous prompt optimization through feedback-guided gradient descent in embedding space.
Few-Shot Text Generation with Natural Language Instructions (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to text generation combine task descriptions and examples with supervised learning.
Approach: They propose a method for text generation that is based on pattern-exploiting training.
Outcome: The proposed approach improves on several summarization and headline generation datasets.
Building Real-World Meeting Summarization Systems using Large Language Models: A Practical Perspective (2023.emnlp-industry)

Copied to clipboard

Challenge: a study examines how to build meeting summarization systems using large language models . closed-source models are generally better in terms of performance, but open-source ones are more advantageous for industrial use .
Approach: They compare closed-source and open-source meeting summarization models for real-world use . they find that closed-sourced models are generally better in terms of performance . however, smaller open-sourced LLMs could still achieve comparable performance if they are open .
Outcome: The proposed model is more efficient for industrial use than closed-source models due to privacy concerns and high cost.
Zero-Shot Strategies for Length-Controllable Summarization (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models struggle with precise length control, particularly in zero-shot settings.
Approach: They propose to use length approximation, target adjustment, sample filtering and automated revisions to improve LLMs' length control capabilities.
Outcome: The proposed methods improve length control in large language models while maintaining or enhancing summary quality without the need for model fine-tuning or architectural changes.
An Exploration of Post-Editing Effectiveness in Text Summarization (2022.naacl-main)

Copied to clipboard

Challenge: Automated summarization methods are efficient but can suffer from low quality.
Approach: They conducted an experiment with 72 participants to compare post-editing provided summaries with manual summarization for summary quality, human efficiency, and user experience.
Outcome: The results show that post-editing improves summary quality, human efficiency, and user experience on formal (XSum news) and informal (Reddit posts) text.
Document Summarization with Latent Queries (2022.tacl-1)

Copied to clipboard

Challenge: Existing benchmarks for query-focused summarization are small for training large neural models.
Approach: They propose a unified modeling framework for query-focused summarization . they model queries as discrete latent variables over document tokens .
Outcome: The proposed framework outperforms strong comparison systems across benchmarks, query types, document settings, and target domains.
DocLens: Multi-aspect Fine-grained Medical Text Evaluation (2024.acl-long)

Copied to clipboard

Challenge: Medical text generation systems are widely used to assist with administrative work and highlight salient information to support decision-making.
Approach: They propose a set of metrics to evaluate completeness, conciseness, and attribution of medical text at a fine-grained level.
Outcome: The proposed framework exhibits substantially higher agreement with medical experts than existing metrics.
Extractive Entity-Centric Summarization as Sentence Selection using Bi-Encoders (2022.aacl-short)

Copied to clipboard

Challenge: Entity-centric summarization is a type of controllable summarizing that aims to produce a summary specific to a given target entity.
Approach: They propose to recast a sentence selection task as a controllable summarization using a dataset supported by EntSUM.
Outcome: The proposed framework outperforms the current state-of-the-art in the sentence selection task and outperformed the competitive entity-centric Lead 3 heuristic by 1.1 F1.
NSTM: Real-Time Query-Driven News Overview Composition at Bloomberg (2020.acl-demos)

Copied to clipboard

Challenge: aggregators consume millions of articles every day, making it difficult to quickly identify key events and miss less-reported stories.
Approach: a new kind of summarization engine was needed to condense large volumes of news into short, easy to absorb points.
Outcome: NSTM can be used to summarize news articles in seconds and quickly and efficiently.
Lessons from the Field: An Adaptable Lifecycle Approach to Applied Dialogue Summarization (2026.eacl-industry)

Copied to clipboard

Challenge: Summarization of multi-party dialogues is a critical capability in industry . but generating high-quality summaries in practice is challenging . prior work has focused on static datasets and benchmarks, a condition rare in practical scenarios .
Approach: They present an agentic system to summarize multi-party interactions using static datasets.
Outcome: The proposed system can summarize multi-party interactions using a set of complex requirements.
Incorporating Question Answering-Based Signals into Abstractive Summarization via Salient Span Selection (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods for generating summarizations using QA-based supervision produce higher quality summaries than baseline methods.
Approach: They propose a method for incorporating question-answering signals into a summarization model by automatically marking document NPs as salient based on whether they are answered in the gold summaries.
Outcome: The proposed method generates higher-quality summaries than baseline methods on benchmark summarization datasets.
SAPGraph: Structure-aware Extractive Summarization for Scientific Papers with Heterogeneous Graph (2022.aacl-main)

Copied to clipboard

Challenge: Abstractive and extractive methods are used to condense long text into concise summaries while retaining essential information.
Approach: They propose to use paper structure to extract paper summaries from long text . they provide a large-scale dataset of COVID-19-related papers .
Outcome: The proposed framework generates more comprehensive and valuable summaries compared to previous work on COVID-19-related papers.
Harmful Factuality: LLMs Correcting What They Shouldn’t (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are trained for factual accuracy, but can conflict with the critical demand for source fidelity.
Approach: They propose a reproducible framework to elicit and measure HFH using controlled entity-level perturbations and strategic entity selection.
Outcome: The proposed framework reduces HFH rates by 50% across summarization, rephrasing, and QA tasks.
Welcome to the Real World: Efficient, Incremental and Scalable Key Point Analysis (2023.emnlp-industry)

Copied to clipboard

Challenge: Key Point Analysis (KPA) extracts the main points from opinions and quantifies their prevalence.
Approach: They propose a key point analysis framework that extracts the main points from opinions and quantifies their prevalence.
Outcome: The proposed system is able to match sentences to key points over five datasets and demonstrate its performance.
LexGenie: Automated Generation of Structured Reports for European Court of Human Rights Case Law (2025.acl-industry)

Copied to clipboard

Challenge: Recent efforts focus on automatic summarization of individual cases, which condense the content of a single case, making it easier for legal professionals to grasp key points.
Approach: They propose a pipeline to generate multi-case structured reports using entire body of case law on user-specified topics within the European Court of Human Rights.
Outcome: The proposed pipeline generates structured reports that enhance efficient, scalable legal analysis.
Abstractive Meeting Summarization: A Survey (2023.tacl-1)

Copied to clipboard

Challenge: Recent advances in deep learning have improved language generation systems, opening the door to improved forms of abstractive summarization.
Approach: They propose to use neural encoder-decoder architectures to generate abstractive meeting summarizations that are particularly well-suited for multi-party conversation.
Outcome: The proposed system could be used in a wide variety of real-world contexts, from business meetings to medical consultations to customer service calls.
Interactive Text Ranking with Bayesian Optimization: A Case Study on Community QA and Summarization (2020.tacl-1)

Copied to clipboard

Challenge: Existing methods that focus on learning a ranking across the whole candidate space are lacking user or task-specific training data.
Approach: They propose an interactive ranking approach that actively selects pairs of candidates, from which the user selects the best.
Outcome: The proposed approach outperforms existing methods in community question answering and extractive multidocument summarization and is an effective reward function for reinforcement learning.
Investigating the Role and Impact of Disfluency on Summarization (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing studies have focused on disfluency detection and removal, with limited studies into its impact on downstream tasks.
Approach: They propose to incorporate disfluency in summarization models to reduce the impact of replacement disfluencies on natural language processing tasks.
Outcome: The proposed model improves on both public and real-life datasets and shows that it can handle disfluent data with up to 6.99-point degradation in Rouge-L score and replacement disfluencies have the highest negative impact.
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards (2025.emnlp-industry)

Copied to clipboard

Challenge: Large language models (LLMs) excel in various tasks, but often produce hallucinations . retrieved contexts, misrepresent information, or generate outright contradictions .
Approach: They propose a framework that measures hallucination faithfulness of large language models . they introduce a leaderboard that leverages diverse human-annotated hallucinian examples .
Outcome: The proposed framework improves hallucination evaluations by leveraging human-annotated examples.
Extending Multi-Document Summarization Evaluation to the Interactive Setting (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches to interactive summarization are incomparable and divergent . a key gap in the development and adoption of interactive summaries is the lack of evaluation methodologies and benchmarks for meaningful comparison of systems.
Approach: They propose an end-to-end evaluation framework for interactive summarization based on expansion-based interaction . framework includes procedure of collecting real user sessions, evaluation measures relying on summarizing standards, but adapted to reflect interaction.
Outcome: The proposed evaluation framework is based on evaluations of baseline implementations and is available publicly as a benchmark.
Improving Zero and Few-Shot Abstractive Summarization with Intermediate Fine-tuning and Data Augmentation (2021.naacl-main)

Copied to clipboard

Challenge: Recent advances in summarization are driven by the availability of large datasets such as the CNN-DailyMail corpus and the New York Times corpus.
Approach: They propose a method for fine-tuning pretrained models for summarization in unsupervised manner . they use Wikipedia data to produce pseudo-summaries which contain characteristics of target dataset .
Outcome: The proposed method achieves state-of-the-art, zero-shot abstractive summarization performance on CNN-DailyMail dataset and compares with other methods on other datasets.
DelucionQA: Detecting Hallucinations in Domain-specific Question Answering (2023.findings-emnlp)

Copied to clipboard

Challenge: Hallucination is a well-known phenomenon in text generated by large language models . state-of-the-art LLMs still have a number of weaknesses, including the tendency to generate hallucinatory statements without considering the factuality .
Approach: They propose a dataset that captures hallucinations made by retrieval-augmented LLMs . they propose to use these methods to help detect hallucinosity in QA tasks .
Outcome: The proposed method captures hallucinations made by retrieval-augmented LLMs for QA tasks.
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews (2026.eacl-industry)

Copied to clipboard

Challenge: Existing approaches to extract actionable suggestions from customer reviews are often mixed-intent, unstructured text.
Approach: They propose a hybrid pipeline that uses a RoBERTa classifier and a precision–recall surrogate to extract actionable suggestions from customer reviews.
Outcome: The proposed pipeline outperforms prompt-only, rule-based, and classifier-only baselines in extraction accuracy and cluster coherence.
Fast Abstractive Summarization with Reinforce-Selected Sentence Rewriting (P18-1)

Copied to clipboard

Challenge: Empirically, we achieve the new state-of-the-art on all metrics (including human evaluation) on the CNN/Daily Mail dataset, as well as significantly higher abstractiveness scores.
Approach: They propose a sentence-level policy gradient method that bridges computation between two neural networks in a hierarchical way while maintaining language fluency.
Outcome: The proposed model achieves state-of-the-art on all metrics and higher abstractiveness scores on the CNN/Daily Mail dataset and faster training convergence than previous models.
Summarizing Medical Conversations via Identifying Important Utterances (2020.coling-main)

Copied to clipboard

Challenge: Applying natural language processing (NLP) techniques to the medical field is a prevailing trend nowadays and has great potential in many applications, such as key information extraction in medical literature.
Approach: They propose to use a hierarchical encoder-tagger model to generate medical conversation summarization by identifying important utterances.
Outcome: The proposed model outperforms baseline models and models and adds conversation-related features to improve performance.
Soft Layer-Specific Multi-Task Summarization with Entailment and Question Generation (P18-1)

Copied to clipboard

Challenge: Recent advances on abstractive summarization have allowed substantial improvements in the quality of the model, but there is still scope for improvement.
Approach: They propose novel multi-task architectures with high-level layer-specific sharing across multiple encoder and decoder layers of the three tasks and soft-sharing mechanisms.
Outcome: The proposed model improves on the CNN/DailyMail and Gigaword datasets and on the DUC-2002 transfer setup.
Efficient Vocabulary Reduction for Small Language Models (2025.coling-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have high computational costs and energy consumption, making their deployment in industrial settings difficult.
Approach: They propose a small language model that compresses the embedding layer and reduces model size without significant loss of performance.
Outcome: The proposed model reduces the embedding layer while maintaining performance while improving accuracy and performance.
Newsroom: A Dataset of 1.3 Million Summaries with Diverse Extractive Strategies (N18-1)

Copied to clipboard

Challenge: a dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications . identifying large, high-quality resources for summarization has called for creative solutions in the past.
Approach: They present a summarization dataset of 1.3 million articles and summaries written by newsrooms of 38 major news publications.
Outcome: The summarization dataset shows high diversity of summarizing styles . authors train existing methods on the data to evaluate its utility and challenges.
HaRiM+: Evaluating Summary Quality with Hallucination Risk (2022.aacl-main)

Copied to clipboard

Challenge: Existing summarization models are limited in measuring the factual inconsistency of generated summaries.
Approach: They propose a decoder overconfidence-regularizing objective as a hallucination risk measurement to better estimate the quality of generated summaries.
Outcome: The proposed metric is reference-free and requires no training or modules . it records state-of-the-art correlation to human judgment on three sets of summary-quality annotations.
Factual Relation Discrimination for Factuality-oriented Abstractive Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing factuality-oriented abstractive summarization models only consider the integration of factual information and ignore the causes of factuual errors.
Approach: They propose a factuality-oriented abstractive summarization model that can identify the causes of factual errors.
Outcome: The proposed model outperforms state-of-the-art models in factual metrics.
Template-based Abstractive Microblog Opinion Summarization (2022.tacl-1)

Copied to clipboard

Challenge: Existing work on Twitter uses extractive summarization to filter through information, but this approach often includes incomplete or redundant information.
Approach: They propose to use Twitter data to generate 3100 gold-standard opinion summaries.
Outcome: The proposed method outperforms previous work on extractive summarization models and fine-tunes to improve performance.
Controllable Summarization with Constrained Markov Decision Process (2021.tacl-1)

Copied to clipboard

Challenge: Existing controllable summarization models do not allow users to specify their preference for a particular attribute of the generated summaries.
Approach: They propose a novel training framework based on Constrained Markov Decision Process (CMDP) that includes a reward function and constraints to facilitate better summarization control.
Outcome: The proposed model can be applied to control important attributes of summarization, including length, covered entities, and abstractiveness, while complying with a given attribute’s requirement.
Token-Level Self-Evolution Training for Sequence-to-Sequence Learning (2023.acl-short)

Copied to clipboard

Challenge: Adaptive training approaches do not consider the variation of learning difficulty in different training steps, making the learning deterministic and sub-optimal.
Approach: They propose a dynamic token-level self-evolution training method that reweighs the training losses of different target tokens based on priors.
Outcome: Empirically, the proposed method yields significant improvements on three translation tasks.
Too Late to Train, Too Early To Use? A Study on Necessity and Viability of Low-Resource Bengali LLMs (2025.coling-main)

Copied to clipboard

Challenge: a new generation of English-oriented Large Language Models significantly outperforms older LLMs on low-resource languages.
Approach: They compare Bengali-oriented LLMs with open-weight and closed-source LLM models . they conclude that there is a need for a Bengali model, but lacks high-quality pretraining data .
Outcome: The proposed model outperforms existing models on Bengali on low-resource languages . the results highlight biases in machine-translated datasets used for Bengali NLP tasks .
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment (2024.emnlp-main)

Copied to clipboard

Challenge: Multilingual human preference data are difficult to obtain at scale, making it challenging to extend this framework to diverse languages.
Approach: They propose a method where a reward model is trained on preference data in one source language and applied to other target languages.
Outcome: The proposed approach is effective under comprehensive evaluation settings, including human evaluation.
Unveiling the Achilles’ Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have highlighted various neural metrics that align well with human evaluations.
Approach: They propose a black-box adversarial framework that generates strong disagreements between human and victim evaluators.
Outcome: The proposed framework can significantly improve the performance of human and victim evaluators.
Unsupervised Extractive Opinion Summarization Using Sparse Coding (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for opinion summarization rely on human annotations, which may not be feasible.
Approach: They propose to perform opinion summarization in an unsupervised manner by using a dictionary learning algorithm that implicitly captures semantic information from the review text.
Outcome: The proposed algorithm performs well on SPACE and AMAZON datasets and performs controllable summarization to generate aspect-specific summaries using only a few samples.
Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing LLMs require a new call to the inference endpoint/API for each new query . repeated calls to the endpoints/AP Is expensive and impractical for many real-world use cases.
Approach: They compare the performance of various LLMs for query-based meeting summarization . they find that combining queries for the same context in a single prompt can be used to minimize repeated calls.
Outcome: The proposed approach reduces the number of calls to the inference endpoints/APIs in meeting summarization tasks.
Efficient Out-of-Domain Detection for Sequence to Sequence Models (2023.findings-acl)

Copied to clipboard

Challenge: Sequence-to-sequence (seq2sequ) models are a ubiquitous tool for text generation but they are not suitable for many other tasks.
Approach: They propose to use UE techniques to identify out-of-domain (OOD) inputs where the model is susceptible to errors.
Outcome: The proposed methods outperform heavyweight ensembles on the task of OOD detection.
An Empirical Study of Building a Strong Baseline for Constituency Parsing (P18-2)

Copied to clipboard

Challenge: Sequence-to-sequence models have been used for natural language generation tasks such as machine translation and summarization.
Approach: They propose to build a strong baseline based on general purpose sequence-to-sequence models for constituency parsing.
Outcome: The proposed model outperforms existing models in natural language generation tasks without any explicit task-specific knowledge or architecture of constituent parsing.
A Mixed Hierarchical Attention Based Encoder-Decoder Approach for Standard Table Summarization (N18-2)

Copied to clipboard

Challenge: Structured data summarization involves generation of summaries from structured input data.
Approach: They propose a hierarchical attention-based encoder-decoder model which leverages the structure in addition to the content of the tables.
Outcome: The proposed model improves on the weathergov dataset by 30% over the current state-of-the-art.
Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive Summarization (2022.acl-long)

Copied to clipboard

Challenge: Abstractive summarization systems still suffer from faithfulness errors, authors say . prior work has proposed models that improve faithfulness, but it is unclear whether this improvement comes from an increased level of extractiveness of the outputs.
Approach: They propose a faithfulness-abstractiveness trade-off curve that serves as a control . they also learn a selector to identify the most faithful and abstractive summary for a given document .
Outcome: The proposed model achieves higher faithfulness scores while being abstractive than the baseline system on two datasets.
Pruning Basic Elements for Better Automatic Evaluation of Summaries (N18-2)

Copied to clipboard

Challenge: Summarization studies work on increasing the scores that are given by automatic evaluation measures.
Approach: They propose a simple but highly effective automatic evaluation measure of summarization, pruned Basic Elements.
Outcome: The proposed measure outperforms ROUGE and BE in most cases and achieves highest correlation coefficient in TAC 2011 AESOP task.
Narrative Embedding: Re-Contextualization Through Attention (2021.emnlp-main)

Copied to clipboard

Challenge: a novel approach to narrative event representation uses attention to re-contextualize events across the whole story . a recent study shows that attention is used to attach event semantics to tokens .
Approach: They propose an unsupervised approach to narrative event representation using attention to re-contextualize events across the whole story.
Outcome: The proposed approach achieves state of the art performance on multiple choice and story cloze tasks.
Annotate the Way You Think: An Incremental Note Generation Framework for the Summarization of Medical Conversations (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets for summarization of medical conversations are limited to conversation-summary pairs . a novel annotation framework is proposed to capture the summarizing process via an annotation task .
Approach: They propose an incremental note generation framework that captures the human summarization process via an annotation task by instructing annotators to first incrementally create a draft note and polish it into a reference note.
Outcome: The proposed framework shows that the human summarization process is much more efficient and accurate than the current method.
Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering (2023.acl-long)

Copied to clipboard

Challenge: Among recent NLP research, multi-document processing is gaining increasing attention due to the need to handle and process an increasing amount of textual data and available documents online.
Approach: They propose to pre-train a generic multi-document model from a cross-document question answering pre-training objective by generating salient sentences from one document and challenging it to recover the sentence from which it was generated.
Outcome: The proposed model outperforms zero-shot GPT-3.5 and GPT-4 in multiple document tasks and generates the correct answer and the salient sentence from a salient document.
D2S: Document-to-Slide Generation Via Query-Based Text Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Existing research efforts to automate the document-to-slide generation process face a critical challenge: no publicly available dataset for training and benchmarking.
Approach: They propose a dataset SciDuet that gathers papers and their corresponding slides from recent years’ NLP and ML conferences.
Outcome: The proposed system outperforms state-of-the-art summarization baselines on both automated ROUGE metrics and qualitative human evaluation.
Exploring Data Augmentation for Code Generation Tasks (2023.findings-eacl)

Copied to clipboard

Challenge: Recent advances in natural language processing have impacted how models are trained for programming language tasks.
Approach: They propose to use augmentation methods that yield consistent improvements in code translation and summarization by up to 6.9% and 7.5% respectively.
Outcome: The proposed methods improve translation and summarization by 6.9% and 7.5% respectively.
Prompt Space Optimizing Few-shot Reasoning Success with Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Prompt engineering is an essential technique for enhancing the abilities of large language models (LLMs) by providing explicit and specific instructions.
Approach: They propose a new approach that uses text embeddings to obtain basis vectors by matrix decomposition and constructs a space for representing all prompts.
Outcome: The proposed approach significantly outperforms state-of-the-art prompt paradigms on ten public reasoning benchmarks.
Extractive is not Faithful: An Investigation of Broad Unfaithfulness Problems in Extractive Summarization (2023.acl-long)

Copied to clipboard

Challenge: Abstractive summarization is less prone to unfaithfulness issues than abstractive summaries . but, unfaitfulness problems, i.e., hallucinating new information, are still a problem in extractive summarisation .
Approach: They propose a typology with five types of broad unfaithfulness problems that can appear in extractive summaries, including and beyond not-entailment.
Outcome: The proposed metric shows that it detects unfaithful summaries faster than existing faithfulness evaluation metrics.
LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization (2023.eacl-main)

Copied to clipboard

Challenge: Human evaluation is labor-intensive, expensive to scale, and difficult to design.
Approach: They propose a set of guidelines for human evaluation of faithfulness in long-form summaries that address the following challenges: (1) How can we achieve high inter-annotator agreement on faithfulness scores? (2) How can our annotator minimize workload while maintaining accurate faithfulness?
Outcome: The proposed framework reduces inter-annotator variance in faithfulness scores while minimizing annotator workload while maintaining accuracy.
EDU-level Extractive Summarization with Varying Summary Lengths (2023.findings-eacl)

Copied to clipboard

Challenge: Existing studies on extractive summarization use finer-grained elementary discourse units . few studies exploited finer grained EDUs with little analysis and justification for the extractive unit selection .
Approach: They propose an extractive model with Varying summary lengths that extracts fixed top-k salient sentences from the document as a summary.
Outcome: The proposed model performs better on ROUGE scores than state-of-the-art models.
ThemePro: A Toolkit for the Analysis of Thematic Progression (2020.lrec-1)

Copied to clipboard

Challenge: Thematic progression is relevant to natural language processing applications dealing with discourse structure, argumentation structure, natural language generation, summarization and topic detection.
Approach: They propose a toolkit for automatic analysis of thematic progression using a web interface.
Outcome: ThemePro provides a visualization of the results including syntactic trees, hierarchical thematicity over propositions and thematic progression over whole texts.
MLD-EA: Check and Complete Narrative Coherence by Introducing Emotions and Actions (2025.coling-main)

Copied to clipboard

Challenge: Existing studies focus on summarization and question-answering tasks, but neglect logical coherence within stories.
Approach: They propose a model that leverages large language models to identify narrative gaps and generate coherent sentences that integrate seamlessly with the story’s emotional and logical flow.
Outcome: The proposed model enhances narrative understanding and story generation, highlighting LLMs’ potential as effective logic checkers in story writing with logical coherence and emotional consistency.
DETQUS: Decomposition-Enhanced Transformers for QUery-focused Summarization (2025.naacl-long)

Copied to clipboard

Challenge: Query-focused tabular summarization is an emerging task in table-to-text generation . traditional transformer-based approaches face challenges due to token limitations and the complexity of reasoning over large tables.
Approach: They propose a system that leverages tabular decomposition alongside a fine-tuned encoder-decoder model to improve summarization accuracy.
Outcome: a new system outperforms the state-of-the-art REFACTOR model in a Query-focused tabular summarization task . the proposed system achieves a ROUGE-L score of 0.4437, outperforming the previous state- of-the art model .
NeuroPrune: A Neuro-inspired Topological Sparse Training Algorithm for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Transformer-based Language Models have become ubiquitous in natural language processing due to impressive performance on various tasks.
Approach: They explore how sparsity affects network topology by exploiting mechanisms seen in biological networks . they show that model-agnostic sparsities are performant across diverse NLP tasks .
Outcome: The proposed model-agnostic sparsity approaches are performant and efficient across NLP tasks.
What’s Wrong? Refining Meeting Summaries with LLM Feedback (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for meeting summarization are limited and lack the robustness and context-based accuracy needed to maintain relevance.
Approach: They propose a multi-LLM correction approach for meeting summarization using a two-phase process that mimics the human review process: mistake identification and summary refinement.
Outcome: The proposed approach improves the quality of a given meeting summarization measured by relevance, informativeness, conciseness, and coherence.
TESS: Text-to-Text Self-Conditioned Simplex Diffusion (2024.eacl-long)

Copied to clipboard

Challenge: Existing models for diffusion generation are expensive and discrete, resulting in a large number of diffusion steps to generate text.
Approach: They propose a text diffusion model that is fully non-autoregressive and employs a new form of self-conditioning and applies the diffusion process on the logit simplex space rather than the learned embedding space.
Outcome: The proposed model outperforms state-of-the-art non-autoregressive models, requires fewer diffusion steps with minimal drop in performance, and is competitive with pretrained autoregressive sequence-to-sequence models.
Long Text and Multi-Table Summarization: Dataset and Method (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing document summarization methods focus on the text and filter out the non-textual content. Existing methods cannot meet the requirements of summarizing long text and multiple tables in each report.
Approach: They propose a dataset for automatic document summarization that uses text and tabular data to produce a concise summary covering the input document's salient information.
Outcome: The proposed method can produce a concise summary covering the input document's salient information.
Advancing Precise Outline-Conditioned Text Generation with Task Duality and Explicit Outline Control (2024.eacl-long)

Copied to clipboard

Challenge: Existing studies on outline-conditioned text generation focus on generating text using provided outlines as rough sketches, but lack of clarity and rationality of the rough outlines hampers quality of the generated text.
Approach: They propose a novel task that requires generating stories based on specific, sentence-level outlines.
Outcome: The proposed framework improves the quality of precise outline-conditioned text generation.
MultiHumES: Multilingual Humanitarian Dataset for Extractive Summarization (2021.eacl-main)

Copied to clipboard

Challenge: a new multilingual summarization model is being developed to help humanitarian experts process large amounts of secondary data to derive situational awareness and guide decision-making.
Approach: They propose to use multilingual documents and annotated snippets to improve extraction of secondary data for humanitarian response experts.
Outcome: The proposed model provides multilingual documents with informative snippets that have been annotated by humanitarian analysts over the past four years.
Structure-Infused Copy Mechanisms for Abstractive Summarization (C18-1)

Copied to clipboard

Challenge: Experimental results show that system summaries struggle to preserve syntactic meaning of source texts.
Approach: They propose to incorporate syntactic information from source sentences into abstractive summaries by structure-infused copy mechanisms.
Outcome: The proposed approach compares favorably to state-of-the-art methods.
A Variational Hierarchical Model for Neural Cross-Lingual Summarization (2022.acl-long)

Copied to clipboard

Challenge: Existing studies on cross-lingual summarization focus on pipeline methods or jointly training an end-to-end model through an auxiliary MT or MS objective.
Approach: They propose a hierarchical model for the cross-lingual summarization task . the model is based on the conditional variational auto-encoder .
Outcome: The proposed model generates better cross-lingual summaries than comparison models in the few-shot setting.
Dating Documents using Graph Convolution Networks (P18-1)

Copied to clipboard

Challenge: Existing approaches for document dating assume accurate knowledge of document date, but this is not always available for arbitrary documents from the Web.
Approach: They propose a Graph Convolutional Network (GCN) based document dating approach which exploits syntactic and temporal graph structures of document in a principled way.
Outcome: The proposed approach outperforms state-of-the-art models on real-world datasets by 19% absolute accuracy points.
On Context Utilization in Summarization with Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large language models excel in abstractive summarization tasks, delivering fluent and pertinent summaries.
Approach: They conduct the first comprehensive study on context utilization and position bias in summarization.
Outcome: The proposed benchmark compares two methods to alleviate position bias in summarization tasks.
TriSum: Learning Summarization Ability from Large Language Models with Structured Rationale (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have advanced tasks like text summarization, but their size and computational demands limit their use in resource-constrained and privacy-centric settings.
Approach: They propose a framework for distilling LLMs’ text summarization abilities into a compact, local model using a curriculum learning strategy that evolves from simple to complex tasks.
Outcome: The proposed framework outperforms baseline models on CNN/DailyMail, XSum, and ClinicalTrial, and improves interpretability by providing insights into the summarization rationale.
Provable Fast Greedy Compressive Summarization with Any Monotone Submodular Function (N18-1)

Copied to clipboard

Challenge: Submodular maximization with the greedy algorithm is an effective approach to extractive summarization.
Approach: They propose a submodular maximization method that is 100 to 400 times faster than existing methods for extractive summarization.
Outcome: The proposed method is 100 to 400 times faster than existing method based on integer-linear-programming formulations and achieves 95%-approximation.
Relational Summarization for Corpus Analysis (N18-1)

Copied to clipboard

Challenge: Existing methods for summarizing textual content are often ignored . relationshipal questions are ubiquitous and varied.
Approach: They propose a method which generates a natural language summary of the relationship between two lexical items in a corpus without reference to a knowledge base.
Outcome: The proposed method generates a natural language summary of the relationship between two lexical items in a corpus without reference to a knowledge base.
Summarizing Procedural Text: Data and Approach (2022.findings-emnlp)

Copied to clipboard

Challenge: Procedural text summarization task is a popular task in the NLP field because of its long length and complexity.
Approach: They propose a procedural text summarization task with two granularity . they propose an Entity-State Graph-based Summarizer (ESGS) which aggregates contextual information for each procedure.
Outcome: The proposed model can summarize the entire procedural text or give an overview for each step or both . Experiments on two datasets confirm the proposed model's effectiveness.
Training Dynamics for Text Summarization Models (2022.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models have shown impressive results when fine-tuned on large summarization datasets.
Approach: They analyze the training dynamics for generation models, focusing on summarization . they find that a propensity to copy the input is learned early in the training process .
Outcome: The proposed model learns at different stages of fine-tuning, the authors show . they show that factual errors are learnt in later stages, but not at high-loss tokens .
ARC: Argument Representation and Coverage Analysis for Zero-Shot Long Document Summarization with Instruction Following LLMs (2026.eacl-long)

Copied to clipboard

Challenge: Argument Representation Coverage (ARC) assesses how well summaries preserve salient arguments . despite their fluency, LLMs frequently hallucinate or omit key content .
Approach: They propose an evaluation framework that assesses how well summaries preserve salient arguments . they use argument representation coverage to distinguish between different information types .
Outcome: The proposed framework assesses how well summaries preserve salient arguments . the authors show that LLMs capture some salient roles but omit critical information .
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation (2026.acl-long)

Copied to clipboard

Challenge: a lack of high-quality English privacy policy corpus optimized for legal clarity and readability is limiting translation of privacy policies . 139 privacy policies are often considered "incomprehensible" due to technical jargon, legal language, and convoluted grammatical structures.
Approach: They propose a high-quality English privacy policy corpus annotated by domain experts . they propose APPSI-139 to summarize and interpret privacy policies in English .
Outcome: The proposed framework outperforms large language models in terms of readability and accuracy.
An Empirical Study of Clinical Note Generation from Doctor-Patient Encounters (2023.eacl-main)

Copied to clipboard

Challenge: Medical doctors spend 52 to 102 minutes per day writing clinical notes from patient encounters.
Approach: They propose to use a new dataset to generate automated and manual clinical notes from doctor-patient conversations in a clinical setting.
Outcome: The proposed model could reduce the time spent writing clinical notes from doctor-patient conversations in a clinical setting.
Unifying Human and Statistical Evaluation for Natural Language Generation (N19-1)

Copied to clipboard

Challenge: Human evaluation captures quality but fails to capture diversity . statistical evaluation fails to catch models that plagiarize from training set .
Approach: They propose a framework which evaluates both diversity and quality based on the optimal error rate of predicting whether a sentence is human-generated.
Outcome: The proposed framework evaluates diversity and quality on summarization and chit-chat dialogue.
Accuracy is not enough: Evaluating Personalization in Summarizers (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing accuracy measures cannot evaluate the degree of personalization of summarization models.
Approach: They propose to use a PENS dataset to analyze the degree of personalization of ten different summarization models.
Outcome: The proposed measure can evaluate the degree of personalization of summarization models using the PENS dataset.
Who Taught You That? Tracing Teachers in Model Distillation (2025.findings-acl)

Copied to clipboard

Challenge: Xu et al., 2006, show that model distillation can imbue efficient small language models with task-specific capabilities competitive with expensive teacher LLMs.
Approach: They propose to distill outputs from a large teacher model to a small student model . they propose to use part-of-speech templates as higher-order linguistic features capable of capturing distinctive signals from teacher models that persist in distilled student outputs.
Outcome: The proposed model distillation technique can imbue efficient small language models with task-specific capabilities competitive with (expensive) teacher LLMs.
Does Summary Evaluation Survive Translation to Other Languages? (2022.naacl-main)

Copied to clipboard

Challenge: a quality summarization dataset requires the production and evaluation of summaries by trained humans and machines.
Approach: They translate a summarization dataset in English and compare its performance to seven languages . they explore equivalence testing as an appropriate statistical paradigm for evaluating correlations between human and automated scoring of summaries .
Outcome: The proposed method could be used in seven languages and compares performance across measures.
Mitigating Data Scarceness through Data Synthesis, Augmentation and Curriculum for Abstractive Summarization (2021.findings-emnlp)

Copied to clipboard

Challenge: a new study explores data manipulation techniques for improving abstractive summarization models without the need for any additional data.
Approach: They propose a method of data synthesis with paraphrasing, data augmentation with sample mixing and curriculum learning with new difficulty metrics based on specificity and abstractiveness.
Outcome: The proposed techniques improve abstractive summarization models without additional data . the proposed techniques can be applied in isolation and when combined .
FinMTEB: Finance Massive Text Embedding Benchmark (2025.emnlp-main)

Copied to clipboard

Challenge: Existing text embedding benchmarks for financial domains are inadequately addressing the nuanced requirements of specialized domains like finance.
Approach: They propose a finance-adapted embedding model that outperforms general-purpose models . they also introduce a new model, Fin-E5, which is also open-sourced .
Outcome: The proposed framework outperforms general-purpose models on financial embedding tasks.
Sources of Hallucination by Large Language Models on Inference Tasks (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are claimed to be capable of Natural Language Inference (NLI)
Approach: They propose to use LLMs to probe their behavior using controlled experiments.
Outcome: The proposed models perform significantly worse on NLI test samples which do not conform to these biases than those which do.
Leveraging Language Models for Summarizing Mental State Examinations: A Comprehensive Evaluation and Dataset Release (2025.coling-main)

Copied to clipboard

Challenge: Mental health disorders affect a significant portion of the global population . access to mental health support is limited in developing countries .
Approach: They evaluated a 12-item descriptive MSE questionnaire and five well-known summarization models . they found that language models can generate coherent MSE summaries for doctors .
Outcome: The proposed model can generate coherent summaries from MSEs in a conversational format.
GENDEX: Generative Data Augmentation Strategy Leveraging External Data for Abstractive Dialogue Summarization (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods to summarize text data are limited by the lack of data.
Approach: They propose a method that uses external data to generate synthetic dialogues from short texts containing people and their interpersonal interactions.
Outcome: The proposed method shows robust performance, generalizability, and scalability regardless of complexity of dialogues.
MReD: A Meta-Review Dataset for Structure-Controllable Text Generation (2022.findings-acl)

Copied to clipboard

Challenge: a new text generation dataset is needed to controllable text summarization, but it lacks the domain knowledge.
Approach: They propose to use existing text generation datasets to leverage input and control signals . they propose to annotate each meta-review sentence manually with a control signal .
Outcome: The proposed method can be used to control the structure of a text generation dataset . it can be applied to a variety of tasks, including a task with a large number of meta-review sentences .
Jointly Learning Semantic Parser and Natural Language Generator via Dual Information Maximization (P19-1)

Copied to clipboard

Challenge: Semantic parsing aims to transform natural language utterances into formal meaning representations (MRs) whereas an NL generator achieves the reverse, the two tasks are often studied separately.
Approach: They propose a method of dual information maximization to regularize the learning process by matching the joint distributions of p and q of NLs.
Outcome: The proposed method empirically maximizes the variational lower bounds of expected joint distributions of NL and MRs.
Can you Summarize my learnings? Towards Perspective-based Educational Dialogue Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: Increasing use of virtual tutors has allowed for more efficient, personalized, and interactive AI-based learning experiences.
Approach: They propose a task of Multi-modal Perspective based Dialogue Summarization (MM-PerSumm) that summarizes educational dialogues from three unique perspectives: the Student, the Tutor, and a Generic viewpoint.
Outcome: The proposed model can summarize educational dialogues from three perspectives, while student-oriented summaries should distill learning points, track progress, and suggest scope for improvement.
AD3: Attentive Deep Document Dater (D18-1)

Copied to clipboard

Challenge: Existing methods to predict creation time of documents are based on time-stamp metadata, but none are available.
Approach: They propose an attention-based neural document dating system which utilizes both context and temporal information in documents in a flexible and principled manner.
Outcome: The proposed system outperforms neural and non-neural baselines on multiple real-world datasets.
Extractive Summarization via ChatGPT for Faithful Summary Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Abstractive summarization methods struggle with generating ungrammatical or even nonfactual contents.
Approach: They evaluate ChatGPT's performance on extractive summarization and compare it with traditional fine-tuning methods on benchmark datasets.
Outcome: The proposed pipeline performs better than abstractive methods on summary faithfulness and in-context learning.
TermDiffuSum: A Term-guided Diffusion Model for Extractive Summarization of Legal Documents (2025.coling-main)

Copied to clipboard

Challenge: Recent studies have explored diffusion models for extractive summarization task, showcasing their remarkable capabilities.
Approach: They propose a term-guided diffusion model for extractive summarization of legal documents that incorporates legal terminology into the model via a well-designed multifactor fusion noise weighting schedule.
Outcome: The proposed model outperforms existing models on a self-constructed legal summarization dataset and achieves improvements of 3.10, 2.84, and 2.89 on three public datasets.
DeModify: A Dataset for Analyzing Contextual Constraints on Modifier Deletion (L18-1)

Copied to clipboard

Challenge: a text fragment is discarded when it has a smaller context, causing it to acquire a new meaning or even become false.
Approach: They build a dataset to study the effect of modifiers on the larger context . they focus on single-word modifiers, the smallest unit that can be considered disposable .
Outcome: The proposed dataset aims to determine whether modifiers can be removed without undesirable consequences.
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Recent work has shown that small, context-dependent shifts in word distributions can be used to apply and detect watermarks, but little work has analyzed the impact of these perturbations on the quality of generated texts.
Approach: They propose a framework that allows for analysis of the impact of watermark settings on the quality of generated texts.
Outcome: The proposed framework provides easy visualization of the quality-detection trade-off of watermark settings.
Unsupervised Abstractive Summarization of Bengali Text Documents (2021.eacl-main)

Copied to clipboard

Challenge: Abstractive summarization systems are difficult to perform due to the unavailability of the parallel data for low-resource languages like Bengali.
Approach: They propose a graph-based unsupervised abstractive summarization system in Bengali text documents that requires only a Part-Of-Speech (POS) tagger and a pre-trained language model trained on Bengali texts.
Outcome: The proposed system outperforms baselines without human-annotated reference summaries on a human-random dataset with Bengali text.
UniSumEval: Towards Unified, Fine-grained, Multi-dimensional Summarization Evaluation for LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for summarization quality evaluation lack diverse input scenarios, focus on narrowly defined dimensions, and struggle with subjective and coarse-grained annotation schemes.
Approach: They propose to use AI to help human annotations and identifie potentially hallucinogenic input texts.
Outcome: The proposed benchmarks improve on existing benchmarks in terms of input diversity, granularity of human annotations, and evaluation dimensions.
Enriching Biomedical Knowledge for Low-resource Language Through Large-scale Translation (2023.eacl-main)

Copied to clipboard

Challenge: Biomedical data and benchmarks are highly valuable but limited in low-resource languages such as English.
Approach: They propose a translation model in Vietnamese that trains a pretrained Encoder-Decoder Transformer model on 20 million translated abstracts.
Outcome: The proposed model can translate and produce both pretrained and supervised biomedical data in two biomedically important domains.
When Do Pre-Training Biases Propagate to Downstream Tasks? A Case Study in Text Summarization (2023.eacl-main)

Copied to clipboard

Challenge: Existing studies have shown that large language models contain linguistic and societal biases, but it is unclear how these biase amplify to downstream tasks.
Approach: They investigate how name-nationality bias propagates from pre-training to downstream tasks . they show that these biases manifest themselves as hallucinations in summarization .
Outcome: The proposed model can reduce the rate of hallucinations, but does not change the types of biases that do appear.
UMSE: Unified Multi-scenario Summarization Evaluation (2023.findings-acl)

Copied to clipboard

Challenge: Summarization quality evaluation is a non-trivial task in text summarization.
Approach: They propose a unified multi-scenario summarization evaluation model that shares cross-sceenario knowledge and uses a self-supervised training paradigm to optimize the model without extra human labeling.
Outcome: The proposed model can achieve comparable performance with existing methods for three evaluation scenarios.
EntSUM: A Data Set for Entity-Centric Extractive Summarization (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for controllable summarization fail to generate entity-centric summaries.
Approach: They propose to use a human-annotated data set EntSUM to generate controllable summarization with a focus on named entities as the aspects to control.
Outcome: The proposed data set shows that existing methods fail to generate entity-centric summaries.
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays (2024.findings-acl)

Copied to clipboard

Challenge: Movie screenplay summarization requires an understanding of long input contexts and elements unique to movies.
Approach: They propose a dataset for movie screenplay summarization that includes movie screenplayers accompanied by their Wikipedia plot summaries.
Outcome: The proposed dataset includes 2200 movie screenplays accompanied by their Wikipedia plot summaries.
Contrastive Aligned Joint Learning for Multilingual Summarization (2021.findings-acl)

Copied to clipboard

Challenge: Existing summarization systems for multilingual text summarizing are limited due to the lack of large-scale data in multiple languages.
Approach: They propose a multilingual summarization system that can understand documents in multiple languages and generate summaries in the corresponding language.
Outcome: The proposed model improves over monolingual models in all languages and transferable to other languages.
Few-shot Query-Focused Summarization with Prefix-Merging (2022.emnlp-main)

Copied to clipboard

Challenge: Query-focused summarization has been considered as an important extension for text summarizing . lack of large-scale datasets hinders its development .
Approach: They propose to integrate text summarization and question answering into a prefix-based pretraining strategy for few-shot learning in query-focused summarizing.
Outcome: The proposed prefix-based pretraining outperforms fine-tuning on query-focused summarization.
Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand (2022.naacl-main)

Copied to clipboard

Challenge: Recent advances on models and metrics should benefit and inform each other, authors argue . bidimensional leaderboards allow for fast, accurate evaluation of language generation models .
Approach: They propose a bidimensional leaderboard that tracks progress in language generation models and metrics for their evaluation.
Outcome: The proposed leaderboards track progress in language generation models and metrics for their evaluation.
Abstractive Summarization of Reddit Posts with Multi-level Memory Networks (N19-1)

Copied to clipboard

Challenge: Abstractive summarization methods suffer from inferior performance compared to extractive methods.
Approach: They propose a reddit TIFU dataset and a new abstractive summarization model . they use multi-level memory networks to store information from different levels of abstraction .
Outcome: The proposed model outperforms state-of-the-art summarization models with multi-level memory . the proposed dataset is highly abstractive and outperformed existing models with the proposed model .
Does Pretraining for Summarization Require Knowledge Transfer? (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing theories claim that pretraining models learn linguistic knowledge from the pretraining corpus, but scientific explanations for these benefits remain unknown.
Approach: They propose to use random character n-grams to test models on real corpora to see if the small residual benefit of using real data could be accounted for by the structure of the pretraining task.
Outcome: The proposed task performs on documents consisting of character n-grams, whereas pretrained models perform on real corpora with no residual benefit.
Not all Hallucinations are Good to Throw Away When it Comes to Legal Abstractive Summarization (2025.naacl-long)

Copied to clipboard

Challenge: Existing models for summarization of legal documents rely on external knowledge to generate abstracts.
Approach: They propose an entity-driven approach that learns the model to generate factual hallucinations . they evaluate legal documents in English and French to evaluate their results .
Outcome: The proposed approach reduces non-factual hallucinations and maximizes summary coverage and factual hallucines at entity-level.
DiscoScore: Evaluating Text Generation with BERT and Discourse Coherence (2023.eacl-main)

Copied to clipboard

Challenge: DiscoScore is a parametrized discourse metric that uses BERT to model discourse coherence . it is weak when operated at system level, and is therefore not reliable in a way to spot improvements .
Approach: They propose a parametrized discourse metric which uses BERT to model discourse coherence from different perspectives.
Outcome: The proposed model outperforms existing models on document-level machine translation and summarization.
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: ChatGPT and GPT-4 are popular as evaluation metric for complex generative tasks . however, they are not ready as human replacements due to significant limitations .
Approach: They conduct extensive analysis to examine the stability and reliability of LLMs as automatic evaluators for abstractive summarization.
Outcome: The proposed methods outperform the commonly used automatic metrics but are not ready for human evaluation due to significant limitations.
From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Recent research in mechanistic interpretability has revealed that Large Language models contain disentangled, human-understandable components.
Approach: They propose a framework that first identifies causal task features through frequency recall and interventional filtering, then selects “Feature-Resonant Data” that maximally activates task features for fine-tuning.
Outcome: The proposed framework outperforms existing models on mathematical reasoning, summarization, and translation tasks while using only 50% of the data.
Domain Aligned Prefix Averaging for Domain Generalization in Abstractive Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies on domain generalization have sophisticated training algorithms.
Approach: They propose a lightweight, weight averaging approach to domain generalization for abstractive summarization using prefix tuning and weight adjusting.
Outcome: The proposed method performs better on four diverse summarization domains compared to baselines.
Modeling Content Importance for Summarization with Pre-trained Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies on content importance do not consider semantics and context when evaluating importance.
Approach: They apply information theory to pre-trained language models to define the concept of importance from the perspective of information amount.
Outcome: Experiments on CNN/Daily Mail and New York Times show that the proposed model can model the importance of content better than previous methods based on F1 and ROUGE scores.
Pre-training for Abstractive Document Summarization by Reinstating Source Text (2020.emnlp-main)

Copied to clipboard

Challenge: Abstractive document summarization models are often trained on limited supervised data . authors present three objectives for pretraining abstractive summarizing models .
Approach: They propose to pre-train a SEQ2SEQ based abstractive summarization model on unlabeled text.
Outcome: The proposed method improves on two benchmark summarization datasets with 19GB of text . the goal is sentence reordering, next sentence generation and masked document generation .
To Point or Not to Point: Understanding How Abstractive Summarizers Paraphrase Text (2021.findings-acl)

Copied to clipboard

Challenge: Abstractive summarization models have seen great improvements in recent years, but there is limited understanding of the strategies different models employ and how they relate their understanding of language.
Approach: They characterize how one popular abstractive model uses an explicit copy/generation switch to control its level of abstraction vs extraction . they find that abstractive summarization models lack the semantic understanding necessary to generate paraphrases that are both abstractive and faithful to the source document.
Outcome: The proposed model uses syntactic boundaries to truncate sentences that are often copied verbatim.
Factual Dialogue Summarization via Learning from Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing models generate fluent and coherent summaries, but inconsistencies can be found in generated summary.
Approach: They propose to use symbolic knowledge distillation to improve the factual consistency of smaller pretrained models for dialogue summarization.
Outcome: The proposed model outperforms baseline models in BART, PEGASUS, and Flan-T5 in factual consistency and accuracy.
LipKey: A Large-Scale News Dataset for Absent Keyphrases Generation and Abstractive Summarization (2022.coling-1)

Copied to clipboard

Challenge: Existing work has addressed each element individually, but this study focuses on LipKey, the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles.
Approach: They propose a novel news dataset that consists of highly absent keyphrases . they combine lips keyphrase and TF-IDF to obtain abstractive summaries .
Outcome: The proposed dataset is the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles.
Literature Retrieval for Precision Medicine with Neural Matching and Faceted Summarization (2020.findings-emnlp)

Copied to clipboard

Challenge: IR for precision medicine often involves looking for multiple pieces of evidence that characterize a patient case.
Approach: They propose a document reranking approach that combines neural query-document matching and text summarization toward such retrieval scenarios.
Outcome: The proposed approach achieves state-of-the-art performance on NIST's TREC-PM track dataset.
Incorporating Distributions of Discourse Structure for Long Document Abstractive Summarization (2023.acl-long)

Copied to clipboard

Challenge: Contemporary leading-edge systems for abstractive (long) text summarization employ Transformer encoderdecoder architectures that only consider the nuclearity annotation .
Approach: They propose to incorporate Rhetorical Structure Theory into a novel summarization model that incorporates both the types and uncertainty of rhetorical relations.
Outcome: The proposed model outperforms state-of-the-art models on automatic metrics and human evaluation.
Mixture Content Selection for Diverse Sequence Generation (D19-1)

Copied to clipboard

Challenge: Generating diverse sequences exhibit semantically one-to-many relationships between source and target sequences.
Approach: They propose to separate diversification from generation using a general plug-and-play module that wraps around and guides an existing encoder-decoder model.
Outcome: The proposed method shows that diversification and generation are separate steps in the same model and that the model is robust.
Twist Decoding: Diverse Generators Guide Each Other (2022.emnlp-main)

Copied to clipboard

Challenge: Using a variety of language generation models, ensembling models is challenging during inference.
Approach: They propose a method that decodes text models that do not assume a shared vocabulary, tokenization or generation order.
Outcome: The proposed method outperforms models decoded in isolation over various scenarios.
Convex Aggregation for Opinion Summarization (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in text autoencoders have significantly improved the quality of the latent space, allowing models to generate consistent text from aggregated latent vectors.
Approach: They develop a framework which searches input-output word overlap for latent vector aggregation.
Outcome: The proposed framework improves the quality of the latent space and establishes state-of-the-art performance on two opinion summarization benchmarks.
HighRES: Highlight-based Reference-less Evaluation of Summarization (P19-1)

Copied to clipboard

Challenge: Existing methods for summarizing documents are inconsistent due to the difficulty of manual evaluation.
Approach: They propose a method where summaries are evaluated by multiple annotators against the source document via manually highlighted salient content.
Outcome: The proposed method improves inter-annotator agreement while highlighting differences among systems.
AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient Content (2024.naacl-long)

Copied to clipboard

Challenge: Existing solutions focus on efficient attentions or divide-and-conquer strategies, but these methods sacrifice global context, leading to incoherent and uninformative summaries.
Approach: They propose to leverage the memory-efficient nature of divide-and-conquer methods while preserving global context.
Outcome: The proposed framework improves informativeness, faithfulness, and coherence over baselines on government reports, meeting transcripts, screenplays, scientific papers, and novels.
Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in abstractive summarization systems produce factually inconsistent text . this is emphasized in tasks like summarizing, which often produce inconsistent text with no input article .
Approach: They use reinforcement learning to optimize for factual consistency and explore trade-offs . they use textual-entailment rewards to optimize the accuracy of the generated summaries .
Outcome: The proposed method improves faithfulness, salience and conciseness of the generated summaries.
CoCoA: Confidence- and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing contrastive decoding methods that handle conflict lack adaptability and can degrade performance in low conflict settings.
Approach: They propose a token-level algorithm for principled conflict resolution and enhanced faithfulness that resolves conflict by utilizing confidence-aware measures and the generalized divergence between parametric and contextual distributions.
Outcome: The proposed algorithm achieves 9.2 points on average in QA, summarization, and long-form question answering (LFQA) benchmarks and improves factuality by 2.5 points on the key benchmarks.
MedicalSum: A Guided Clinical Abstractive Summarization Model for Generating Medical Reports from Patient-Doctor Conversations (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing models for summarizing medical conversations do not take clinical knowledge into account and are difficult to control.
Approach: They propose a transformer-based sequence-to-sequence architecture for summarizing medical conversations by integrating medical domain knowledge from the Unified Medical Language System (UMLS).
Outcome: The proposed model achieves state-of-the-art ROUGE score improvements of 0.8-2.1 points (including 6.2% error reduction in the PE section) it incorporates medical domain knowledge from the Unified Medical Language System (UMLS).
Prefix-Tuning: Optimizing Continuous Prompts for Generation (2021.acl-long)

Copied to clipboard

Challenge: Fine-tuning is the prevalent paradigm for using large pretrained language models for downstream tasks, but it requires updating and storing all the parameters of the LM.
Approach: They propose a lightweight alternative to fine-tuning for natural language generation tasks that optimizes a sequence of continuous vectors, which they call the prefix.
Outcome: The proposed approach outperforms fine-tuning in the full data setting and extrapolates better to examples with topics that are unseen during training.
Bias in News Summarization: Measures, Pitfalls and Corpora (2024.findings-acl)

Copied to clipboard

Challenge: Pretrained large language models can reproduce harmful social biases in constrained settings, such as summarization.
Approach: They propose a method to generate input documents with carefully controlled demographic attributes and then apply it to a controlled setting.
Outcome: The proposed method allows to generate input documents with carefully controlled demographic attributes while working with real-world input documents.
Multi-Reference Training with Pseudo-References for Neural Translation and Text Generation (D18-1)

Copied to clipboard

Challenge: Neural text generation has been quite successful recently, but during training time, only one reference is considered for each example, even though there are often multiple references available.
Approach: They propose an algorithm to generate exponentially many pseudo-references by compressing existing references into lattices and traversing them to generate new pseudo-References.
Outcome: The proposed model significantly improves on baselines in machine translation and image captioning, and is comparable to existing models.
PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document Summarization (2022.acl-long)

Copied to clipboard

Challenge: Existing pretrained models require domain-specific additional information to be effective.
Approach: They propose a pre-trained model for multi-document representation with a focus on summarization that uses efficient encoder-decoder transformers to simplify the processing of concatenated input documents.
Outcome: PRIMERA outperforms current state-of-the-art models on most datasets with large margins . PRImerA uses efficient encoder-decoder transformers to simplify processing of concatenated input documents.
ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Detecting contradictions in texts is often regarded as determining relation between hypothesis and piece of premise.
Approach: They propose a human-annotated dataset to study self-contradictions in long documents . they analyze the capabilities of four open-source and commercially available LLMs .
Outcome: The proposed dataset outperforms open-source LLMs on document-level tasks but struggles with self-contradictions that require more nuance and context.
ACLSum: A New Dataset for Aspect-based Summarization of Scientific Publications (2024.naacl-long)

Copied to clipboard

Challenge: Existing statistical phrasal or hierarchical machine translation systems relies on a large set of translation rules which results in engineering challenges.
Approach: They propose to use factorized grammar from the field of linguistics as more general translation rules from XTAG English Grammar to generate a manually crafted summarization dataset.
Outcome: The proposed method outperforms existing methods on low-resource language translation tasks with less training data.
An Exploratory Study on Long Dialogue Summarization: What Works and What’s Next (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing models for dialogue summarization focus on extracting the main events of short conversations, but real-world dialogues are difficult to train.
Approach: They propose three strategies to deal with the lengthy input problem and locate relevant information using long dialogue datasets.
Outcome: The retrieve-then-summarize pipeline models yield the best performance on three long dialogue datasets.
PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have raised concerns about the potential threats large language models pose to academic integrity and copyright protection.
Approach: They propose a dataset of 46.5K synthetic text pairs that represent three major types of plagiarism: verbatim copying, paraphrasing, and summarization.
Outcome: The proposed dataset shows that GPT-3.5 Turbo can produce high-quality paraphrases and summaries without significantly increasing text complexity compared to GPT-4 Turbo.
VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions (2025.acl-long)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) can grasp the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent.
Approach: They propose a framework for multimodal large language models to grasp the intention of a question and decompose it into a series of visual recognition sub-tasks to find out the answer.
Outcome: The proposed framework improves the accuracy of complex video-related questions by 29.6% and 17.2% on CVQA and the existing VQA datasets.
Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy (2021.emnlp-main)

Copied to clipboard

Challenge: Statistical language modeling and translation with transformers have found many successful applications in program understanding and generation tasks.
Approach: They propose an architecture-independent approach for leveraging syntactic hierarchies of source code . they use syntax trees to extract syntak hierarchical structures and integrate them into context window .
Outcome: The proposed approach achieves state-of-the-art in code completion and summarization for Python in the CodeXGLUE benchmark.
How “Multi” is Multi-Document Summarization? (2022.emnlp-main)

Copied to clipboard

Challenge: Multi-document summarization (MDS) aims at combining information spread across multiple documents . a single document often covers the full summary content .
Approach: They propose a measure to evaluate the degree to which a summary is "disperse" they propose to combine information from multiple documents into a single document to generate a concise summary .
Outcome: The proposed measure evaluates the degree to which a summary is "disperse" the measure is applied to several popular MDS datasets and state-of-the-art systems.
ForumSum: A Multi-Speaker Conversation Summarization Dataset (2021.findings-emnlp)

Copied to clipboard

Challenge: Abstractive summarization quality has been improved but there is a lack of data for conversation summarizing applications.
Approach: They propose to build a conversation summarization dataset with human written summaries from internet forums.
Outcome: The proposed dataset can be easily expanded to improve conversation summarization applications.
What’s under the hood: Investigating Automatic Metrics on Meeting Summarization (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation metrics do not capture meeting-specific errors, leading to ineffective assessment.
Approach: They examine the relationship between established metrics and human evaluations to determine what challenges and errors are captured by correlating metric scores with human evaluation.
Outcome: The proposed measures show weak correlations with human evaluations and a third of the correlations show error masking.
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) generate misleading or outright incorrect information.
Approach: They propose a method that debiases uncertainty scores on output length and uses residuals as corrected, length-invariant estimates.
Outcome: The proposed method improves over nominally length-normalized methods on machine translation, summarization, and question-answering tasks.
PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models exhibit positional bias, a problem that can undermine the completeness of conversation summarizations.
Approach: They propose a semantic similarity-based sentence-level metric to quantify positional bias in conversational summaries.
Outcome: The proposed benchmark provides the first systematic evaluation of positional bias in conversational summarization across languages and contexts.
Nutri-bullets Hybrid: Consensual Multi-document Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for generating comparative summaries that highlight similarities and contradictions in input documents are lacking large parallel training data for their training.
Approach: They propose a method for generating comparative summaries that highlight similarities and contradictions in input documents by using a neural interpretation of traditional concept-to-text generation systems.
Outcome: The proposed model is compared with conventional methods in the domain of nutrition and health, where the existing models lack large parallel training data.
Unsupervised Learning of Hierarchical Conversation Structure (2022.findings-emnlp)

Copied to clipboard

Challenge: Goal-oriented conversations often have sub-dialogue structure, but it can be domain-dependent . Increasingly, language understanding applications involve conversational speech and text .
Approach: They propose an unsupervised approach to learning hierarchical conversation structure . they use turn and sub-dialogue segment labels to decode the structure based on dialogue acts and subtasks .
Outcome: The proposed approach improves neural models for three conversation-level understanding tasks.
SciReviewGen: A Large-scale Dataset for Automatic Literature Review Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing literature review models have addressed literature review generation, but lack of large-scale datasets has been a stumbling block.
Approach: They propose to use a large-scale dataset to evaluate automatic literature review generation models.
Outcome: The proposed model can generate summaries comparable to human-written reviews while lacking detailed information.
BioGen: Generating Biography Summary under Table Guidance on Wikipedia (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for summarizing text have not captured the salient information from an article.
Approach: They propose a table-guided abstractive biography summarization that utilizes factual tables to capture important information and generate a summary of a biography.
Outcome: The proposed method is the first large-scale biography summarization dataset with tables.
Warmup Generations: A Task-Agnostic Approach for Guiding Sequence-to-Sequence Learning with Unsupervised Initial State Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing supervised fine-tuning (SFT) methods focus on directly generating the target output without leveraging the benefits of intermediate steps or initial guidance.
Approach: They propose a task-agnostic framework that enables models to generate intermediate "warmup" sequences that are iteratively refined to maximize their contribution to the final output.
Outcome: The proposed framework outperforms traditional supervised fine-tuning methods on translation, summarization, and multi-choice question answering tasks.
DACSA: A large-scale Dataset for Automatic summarization of Catalan and Spanish newspaper Articles (2022.naacl-main)

Copied to clipboard

Challenge: a large corpus of documents is available for summarization tasks in English . supervised methods require adequate corpora for summarizing .
Approach: They describe a corpus of catalan and spanish newspapers that can be used to train summarization models for Catalan, Spanish and other languages.
Outcome: The proposed corpus can be used to train summarization models for Catalan and Spanish.
DocNLI: A Large-scale Dataset for Document-level Natural Language Inference (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on sentence-level inference, which limits its application in downstream NLP problems.
Approach: They propose to construct a large-scale dataset for document-level NLI that can be used to study NLP problems.
Outcome: The proposed model performs well on popular sentence-level benchmarks and generalizes well to out-of-domain NLP tasks that rely on inference at document granularity.
Did You Get It? A Zero-Shot Approach to Locate Information Transfers in Conversations (2024.lrec-main)

Copied to clipboard

Challenge: Existing models do not provide an efficient way to locate information that enters the common ground.
Approach: They propose a method based on segmentation of a conversation into themes followed by their summarization and obtain the location of information transfers by computing the distance between the theme summary and the different utterances produced by a speaker.
Outcome: The proposed method is based on the segmentation of a conversation into themes followed by their summarization and obtains the location of information transfers by computing the distance between the theme summary and the different utterances produced by a speaker.
Iterative Document Representation Learning Towards Summarization with Polishing (D18-1)

Copied to clipboard

Challenge: Existing summarization methods read through document only once to generate a document representation, resulting in a sub-optimal representation.
Approach: They propose an iterative model for supervised extractive text summarization which polishes the document representation on many passes through the document.
Outcome: The proposed model outperforms state-of-the-art extractive systems on CNN/DailyMail and DUC2002 datasets.
Facet-Aware Evaluation for Extractive Summarization (2020.acl-main)

Copied to clipboard

Challenge: lexical overlap is a common evaluation metric for extractive summarization, but recent studies reveal its limitations.
Approach: They propose a facet-aware evaluation setup for better assessment of information coverage in extractive summaries.
Outcome: The proposed evaluation setup improves human correlation with extractive summarization datasets and improves comparative analysis.
ASPECTNEWS: Aspect-Oriented Summarization of News Documents (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for generating generic summarizations can't be used to generalize to these domains without seeing in-domain training data.
Approach: They use a dataset of real-world aspect-oriented summaries to annotate articles from two different news sub-domains.
Outcome: The proposed approach produces better focused summaries than existing systems without seeing in-domain training data.
Hierarchical Transformer for Task Oriented Dialog Systems (2021.naacl-main)

Copied to clipboard

Challenge: Existing models for dialog generation are challenging to train using the standard Seq2Seq models.
Approach: They propose a framework for Hierarchical Transformer Encoders that can be morphed into any hierarchical transformer by using specially designed attention masks and positional encodings.
Outcome: The proposed framework can be morphed into any hierarchical encoder, including HRED and HIBERT like models, by using specially designed attention masks and positional encodings.
Discourse-Aware Neural Extractive Text Summarization (2020.acl-main)

Copied to clipboard

Challenge: Recent studies have shown that sentence-based extractive models result in redundant or uninformative phrases in the extracted summaries.
Approach: They propose a discourse-aware neural summarization model that extracts sub-sentential discourse units as candidates for extractive selection on a finer granularity.
Outcome: Experiments show that the proposed model outperforms state-of-the-art models on popular summarization benchmarks.
Discrete Optimization for Unsupervised Sentence Summarization with Word-Level Extraction (2020.acl-main)

Copied to clipboard

Challenge: Sentence summarization systems that use latent space to reconstruct the source sentence are unwillingly exploited.
Approach: They propose a method that uses language modeling and semantic similarity metrics to find a high-scoring summary.
Outcome: The proposed method achieves state-of-the-art for unsupervised sentence summarization according to ROUGE scores.
Distinguishing Between Foreground and Background Events in News (2020.coling-main)

Copied to clipboard

Challenge: a new task is needed to distinguish between foreground and background events in news articles .
Approach: They propose a task of distinguishing between foreground and background events in news articles . they also identify the general temporal position of background events relative to the foregoing period .
Outcome: The proposed model achieves good performance on a dataset of news articles .
Distribution Aware Metrics for Conditional Natural Language Generation (2024.lrec-main)

Copied to clipboard

Challenge: Existing metrics for conditional natural language generation rely on pairwise comparisons between a single generated text and the best-matching reference.
Approach: They propose a family of meta-metrics that build on existing pairwise distance functions to evaluate conditional natural language generation models.
Outcome: The proposed method evaluates the ability of a model to generate text matching diversity in references in visual description and summarization.
FACTS: Table Summarization via Offline Template Generation with Agentic Workflows (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for query-focused table summarization struggle with complex reasoning and token-limit issues.
Approach: They propose a Fast, Accurate, and Privacy-Compliant table summarization approach via Offline Template Generation.
Outcome: The proposed method outperforms baseline methods on widely-used benchmarks.
Hierarchical Catalogue Generation for Literature Review: A Benchmark (2023.findings-emnlp)

Copied to clipboard

Challenge: Scientific literature review generation aims to extract and organize important information from an abundant collection of reference papers and produces corresponding reviews while lacking a clear and logical hierarchy.
Approach: They propose a task to generate a hierarchical catalogue of a review paper given various references by using a database of 7.6k literature review catalogues and 389k reference papers.
Outcome: The proposed method produces a hierarchical catalogue of a review paper given various references.
Enhancing Argument Summarization: Prioritizing Exhaustiveness in Key Point Generation and Introducing an Automatic Coverage Evaluation Metric (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for summarizing arguments are incapable of distinguishing between generated key points of different qualities.
Approach: They propose an extractive approach that generates concise, high quality key points . they propose to use a clustering approach to generate key points from raw arguments .
Outcome: The proposed method outperforms state-of-the-art methods for key point generation . it offers concise, high quality generated key points with higher coverage of reference summaries .
Hooks in the Headline: Learning to Generate Headlines with Controlled Styles (2020.acl-main)

Copied to clipboard

Challenge: Current summarization systems only produce plain, factual headlines, far from the practical needs for exposure and memorableness of the articles.
Approach: They propose a task to generate relevant headlines with three style options . they propose combining summarization and reconstruction tasks into a multitasking framework .
Outcome: The proposed method outperforms the state-of-the-art summarization model by 9.68% . it can generate relevant, fluent headlines with humor, romance and clickbait .
Gradient-guided Attention Map Editing: Towards Efficient Contextual Hallucination Mitigation (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) often experience “contextual hallucination” where they prioritize self-generated content over input context, leading to a disregard for pertinent details.
Approach: They propose a method that dynamically adjusts attention maps to enhance contextual relevance by using a trained classifier to identify attention maps likely to induce hallucinations.
Outcome: The proposed approach reduces hallucinations across open-source models on summarization and open-book QA tasks.
Multi-Stage Pre-training Enhanced by ChatGPT for Multi-Scenario Multi-Domain Dialogue Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for dialogue summarization only apply to specific scenarios and domains.
Approach: They propose a pre-trained model specifically designed for multi-scenario multi-domain dialogue summarization.
Outcome: The proposed model significantly outperforms state-of-the-art models on three dialogue summarization datasets from different scenarios and domains.
Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack (2024.acl-long)

Copied to clipboard

Challenge: Recent developments in balancing usefulness and safety of large language models raise a critical question . current attacks, especially adversarial ones that manipulate malicious prompts, often aim to manipulate the input .
Approach: They show that LLMs can effectively summarize malicious long documents but often refuse to translate them.
Outcome: The findings highlight a vulnerability in LLMs that can't translate or summarize documents . the study focuses on LLM models, Gemini and GPT-4, which can' be exploited .
Unsupervised Opinion Summarization as Copycat-Review Generation (2020.acl-main)

Copied to clipboard

Challenge: Recent work on opinion summarization has focused on extracting fragments from reviews, but we use novel sentences to generate abstractive summaries.
Approach: They propose an abstractive summarizer which does not use summaries in training and is trained end-to-end on a large collection of reviews.
Outcome: The proposed model produces fluent and coherent summaries reflecting consensus opinions on Amazon and Yelp reviews.
Matina: A Large-Scale 73B Token Persian Text Corpus (2025.naacl-long)

Copied to clipboard

Challenge: Existing Persian datasets are small and lack content diversity . lack of high-quality data has slowed development of NLP models and open-source LLMs for Persian.
Approach: They propose a Persian dataset of 72.9B tokens that is preprocessed and deduplicated to ensure high data quality.
Outcome: The proposed model performs well on key Persian NLP tasks.
Does the Generator Mind Its Contexts? An Analysis of Generative Model Faithfulness under Context Transfer (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have focused on examining hallucinations stemming from static input, such as in summarization or machine translation.
Approach: They propose a knowledge-augmented generator that produces information that remains grounded in contextual knowledge regardless of alterations in the context.
Outcome: The proposed method is designed to produce information that remains grounded in contextual knowledge, regardless of alterations in the context.
A Mixed-Language Multi-Document News Summarization Dataset and a Graphs-Based Extract-Generate Model (2025.naacl-long)

Copied to clipboard

Challenge: Existing research on news summarization focuses on single-language single-document (SLSD), single-linguistic multi-document or cross-language multi-doc (CLSD) however, in real-world scenarios, news articles often involve multiple documents in different languages, i.e., mixed-language MLMD.
Approach: They propose a mixed-language multi-document news summarization dataset with four different languages and 10,992 source document cluster and target summary pairs.
Outcome: The proposed dataset contains four different languages and 10,992 source document cluster and target summary pairs.
OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory (2026.acl-long)

Copied to clipboard

Challenge: Existing LLMs are limited by text-context budgets, resulting in token-expensive storage of raw trajectories . Optical Context Retrieval Memory (OCR-Memory) renders historical tra-jectorios into images annotated with unique visual identifiers.
Approach: They propose a framework that leverages the visual modality as a high-density representation of agent experience.
Outcome: Optical Context Retrieval Memory (OCRM) renders historical trajectories into images annotated with unique visual identifiers.
Improving Faithfulness in Abstractive Summarization with Contrast Candidate Generation and Selection (2021.naacl-main)

Copied to clipboard

Challenge: Recent studies have shown that current models are prone to generating unfaithful summaries . a proposed method is effective in identifying and correcting extrinsic hallucinations .
Approach: They propose a model-agnostic post-processing technique to correct unfaithful summaries . they generate alternative candidates where names and quantities are replaced with compatible ones .
Outcome: The proposed method corrects extrinsic hallucinations in unfaithful summaries.
CFSum Coarse-to-Fine Contribution Network for Multimodal Summarization (2023.acl-long)

Copied to clipboard

Challenge: Existing multimodal summarization models ignore the contribution of visual modalities . we propose a novel contribution network to consider different contributions of images .
Approach: They propose a Coarse-to-Fine contribution network for multimodal summarization to consider different contributions of images for summarizing.
Outcome: The proposed system outperforms baselines on the visual and textual modalities.
X-FACTOR: A Cross-metric Evaluation of Factual Correctness in Abstractive Summarization (2022.emnlp-main)

Copied to clipboard

Challenge: Abstractive summarization models produce factually inconsistent summaries that are not supported by the original article.
Approach: They propose a fact-aware filtering mechanism that improves the factuality of abstractive summarization models.
Outcome: The proposed method improves the quality of training data and the factuality of generated summaries.
On Learning to Summarize with Large Language Models as References (2024.naacl-long)

Copied to clipboard

Challenge: Recent studies have found that summaries generated by large language models (LLMs) are favored by human annotators when compared to reference summary from widely used summarization datasets.
Approach: They propose to use large language models (LLMs) as reference learning settings for smaller text summarization models to investigate whether their performance can be substantially improved.
Outcome: The proposed model outperforms standard supervised fine-tuning and human evaluations while retaining human-level performance.
Separating Context and Pattern: Learning Disentangled Sentence Representations for Low-Resource Extractive Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Context information is one of the key factors for extractive summarization, but other factors can be used to identify sentence importance.
Approach: They propose to disentangle context and pattern factors for extractive summarization . they separate context and patterns for a better generalization ability in low-resource setting .
Outcome: The proposed model can be used in the zero-shot setting or fine-tuned in the few-shot settings.
Controllable Abstractive Sentence Summarization with Guiding Entities (2020.coling-main)

Copied to clipboard

Challenge: Existing text summarization models lack guiding entities to ensure that entities are present in summaries.
Approach: They propose a controllable abstractive sentence summarization model which generates summaries with guiding entities.
Outcome: The proposed model outperforms the state-of-the-art models in evaluation scores and informativeness metrics.
HOLMS: Alternative Summary Evaluation with Large Language Models (2020.coling-main)

Copied to clipboard

Challenge: Efficient document summarization requires evaluation measures that can rank a set of systems based on an average score and highlight which individual summary is better than another.
Approach: They propose a hybrid evaluation measure for document summarization called HOLMS that combines both language models pre-trained on large corpora and lexical similarity measures.
Outcome: The proposed measure outperforms ROUGE and BLEU on several extractive summarization datasets for both linguistic quality and pyramid scores.
TWEETSUM: Event oriented Social Summarization Dataset (2020.coling-main)

Copied to clipboard

Challenge: Developing social summarization systems is becoming more and more critical . but, the publicly available and high-quality large scale social summaries are rare .
Approach: They propose to build a social summarization dataset using twitter's hot events . they collect user relations, hashtags and user profiles to evaluate their summarizing methods .
Outcome: The proposed dataset is based on a dataset from twitter with 12 real world hot events with 44,034 tweets and 11,240 users.
Background Summarization of Event Timelines (2023.emnlp-main)

Copied to clipboard

Challenge: Generating concise summaries of news events is a challenging task for newcomers to a news story.
Approach: They propose a task of background news summarization that complements each timeline update with a background summary of relevant preceding events.
Outcome: The proposed system performs well on a question-answering-based evaluation metric, Background Utility Score (BUS).
Factual Error Correction for Abstractive Summarization Models (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for abstractive summarization are unable to ensure factual consistency of generated summaries.
Approach: They propose a post-editing corrector module to identify and correct factual errors in generated summaries.
Outcome: The proposed model outperforms existing models on CNN/DailyMail dataset on factual consistency evaluation.
Compressive Summarization with Plausibility and Salience Modeling (2020.emnlp-main)

Copied to clipboard

Challenge: a new method to learn which compressions to apply is based on syntactic rules for deleting spans . plausibility and salience are the two main criteria for determining which compression to apply . a recent study shows that the plausability model generally selects for grammatical and factual deletions compared to extractive methods .
Approach: They propose to leave the decision about what to delete to two data-driven criteria . they show that plausibility and salience are the most important criteria if a span is deleted .
Outcome: The proposed method achieves strong in-domain results on benchmark datasets and human evaluation shows that plausibility model generally selects for grammatical and factual deletions.
Better Highlighting: Creating Sub-Sentence Summary Highlights (2020.emnlp-main)

Copied to clipboard

Challenge: Abstractive summarizations are considered to be less reliable because they distort the original meaning and can be confusing for readers.
Approach: They propose a method to generate summary highlights that are understandable on their own to avoid confusion.
Outcome: The proposed method allows summaries to be understood in context and avoids misdirecting readers to false conclusions.
Enhancing Factual Consistency in Text Summarization via Counterfactual Debiasing (2025.coling-main)

Copied to clipboard

Challenge: Abstractive text summarization has produced fluent and informative outputs, but factual inconsistency is a challenge.
Approach: They propose a framework that mitigates the causal effects of language bias and irrelevancy bias by counterfactual estimation.
Outcome: The proposed framework outperforms baseline methods on two widely used summarization datasets.
CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for generating abstractive summarization are inconsistent and rely on heuristically created data for error handling.
Approach: They propose a contrastive learning formulation that leverages both positive and negative summaries to train summarization systems that are better at distinguishing between them.
Outcome: The proposed learning framework produces more factual summaries than strong comparisons with post error correction, entailment-based reranking, and unlikelihood training.
CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLM (2025.emnlp-main)

Copied to clipboard

Challenge: Experiments show that CIFLEX significantly reduces computational costs without degrading task performance.
Approach: They propose a new execution system for efficient sub-task handling with a single large language model.
Outcome: Experiments show that CIFLEX significantly reduces computational costs without degrading task performance.
Enhancing Scientific Document Summarization with Research Community Perspective and Background Knowledge (2024.lrec-main)

Copied to clipboard

Challenge: Scientific paper summarization is the focus of recent research . prevailing summarizing methods involve selective extraction of content from abstract, introduction, and conclusion segments within the target articles.
Approach: They propose a model that incorporates references and citations to capture the impact of the document on the research community.
Outcome: The proposed model generates extractive and abstractive summaries in parallel and improves their performance when considering the standard metrics.
Phrase-Level Localization of Inconsistency Errors in Summarization by Weak Supervision (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for evaluating inconsistency in summarization are limited . a recent study found that more than 30% of summarized summaries are inconsistent with the source documents .
Approach: They propose a method for localizing inconsistency errors in summarization using a synthetic dataset that contains factual errors likely to be produced by a common language processor.
Outcome: The proposed method detects factual errors more accurately than existing weakly supervised methods . the proposed model also detects errors in original sentences more accurately .
Summarization of Opinionated Political Documents with Varied Perspectives (2025.coling-main)

Copied to clipboard

Challenge: Political ideologies can lead people to develop misperceptions of groups with opposing opinions, such as the 2024 US presidential election, French legislative election, or the Brexit referendum.
Approach: They propose a dataset and task for independently summarizing political perspectives in a set of opinionated news articles.
Outcome: The proposed dataset and task evaluates models of varying sizes and architectures on a set of opinionated news articles.
Dissecting Generation Modes for Abstractive Summarization Models via Ablation and Attribution (2021.acl-long)

Copied to clipboard

Challenge: Abstractive summarization models have made great strides in recent years, but little is known about how they actually form summaries and how to understand where their decisions come from.
Approach: They propose a two-step method to interpret summarization model decisions by categorizing each decoder decision into one of several generation modes.
Outcome: The proposed method can identify phrases the summarization model has memorized and determine where in the training pipeline this memorization happened, and study complex generation phenomena on a per-instance basis.
ArgLegalSumm: Improving Abstractive Summarization of Legal Documents with Argument Mining (2022.coling-1)

Copied to clipboard

Challenge: Existing abstractive summarization models do not take into account argumentative structure of legal documents, which poses a challenge towards effective abstractive summary.
Approach: They propose a technique that integrates argument role labeling into the summarization process by integrating argument role labels into the document.
Outcome: The proposed method improves over strong baselines with pretrained language models.
Predicting Text Preference Via Structured Comparative Reasoning (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches to comparative reasoning rely on pretraining or fine-tuning models at the cost of massive human annotation and computation.
Approach: They propose a model that prompts LLMs to generate structured intermediate comparisons by proposing aspects for comparison, followed by generating textual comparisons under each aspect.
Outcome: The proposed model significantly reduces hallucination and improves consistency across various NLP tasks.
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting (2024.emnlp-main)

Copied to clipboard

Challenge: Pretrained large language models excel in a variety of natural language processing tasks . however, they pose significant security risks due to their tendency to memorize training data .
Approach: They propose a method to estimate LLM memorization using dynamic, prefix-dependent soft prompts.
Outcome: The proposed method can achieve maximum relative improvement of 135.3% and 39.8% over baseline compared to state-of-the-art methods.
PORT: Preference Optimization on Reasoning Traces (2025.naacl-long)

Copied to clipboard

Challenge: Preference optimization methods have been successfully applied to improve the alignment of large language models with human values.
Approach: They propose to use preference optimization methods to generate rejected answers using weak LLM prompting and digit corruption to improve the mathematical reasoning abilities of language models.
Outcome: The proposed method leads to increased accuracy on the GSM8K and AQuA-RAT benchmarks without annotations.
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in efficient attention mechanisms have led to the expansion of the context length of large language models.
Approach: They propose a procedure to synthesize Haystacks of documents and generate a summary that identifies relevant insights and precisely cites the source documents.
Outcome: The proposed evaluation can score summaries on Coverage and Citation . the proposed evaluation lags human performance estimates by 10+ points on SummHay .
Tutor-ICL: Guiding Large Language Models for Improved In-Context Learning Performance (2024.findings-emnlp)

Copied to clipboard

Challenge: In-context learning (ICL) is a dominant paradigm in natural language processing.
Approach: They propose a prompting method for classification tasks using exemplar answers in a *comparative format' they also propose introducing a test instance before the exemplars to improve performance .
Outcome: The proposed method achieves up to 13.76% increase in accuracy on classification tasks across decoder-only and encoder-decoder LLMs.
Segmented Recurrent Transformer: An Efficient Sequence-to-Sequence Model (2023.findings-emnlp)

Copied to clipboard

Challenge: Transformers have shown dominant performance across a range of domains including language and vision, but their computational cost grows quadratically with the sequence length, making their usage prohibitive for resource-constrained applications.
Approach: They propose a segmented recurrent transformer that combines segmente recursion with recursive attention to reduce the computational cost.
Outcome: The proposed model achieves higher ROUGE1 scores and lower computational complexity than current approaches.
TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship Classification (2025.acl-long)

Copied to clipboard

Challenge: Large language models have demonstrated great potential in natural language generation, but their widespread adoption has raised concerns regarding content reliability and accountability.
Approach: They propose a challenge to trace each sentence of a target text back to specific source sentences within potentially lengthy or multi-document inputs.
Outcome: The proposed challenge traces each sentence of a target text back to specific source sentences . the dataset includes 11 scenarios covering QA and summarization in english and Chinese .
Toward Multi-Session Personalized Conversation: A Large-Scale Dataset and Hierarchical Tree Framework for Implicit Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing long-term open-domain dialogue datasets lack complex, real-world personalization and fail to capture implicit reasoning.
Approach: They propose a large-scale long-term dataset with 2,500 examples containing approximately 100 conversation sessions to study implicit reasoning in personalized dialogues.
Outcome: The proposed model improves the ability of LLMs to reason over long-term conversations with implicit contextual dependencies.
AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge (2025.naacl-long)

Copied to clipboard

Challenge: Existing contrastive methods that ignore the context of a large language model (LLM) fail to handle instances that vary in their amount of conflict, with static methods over-adjusting when conflict is absent.
Approach: They propose a fine-grained, instance-level approach called AdaCAD which dynamically adjusts the degree of conflict based on the degree.
Outcome: The proposed approach outperforms baselines and improves factuality of summaries by 6.19.
SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation (2023.emnlp-main)

Copied to clipboard

Challenge: evaluating the quality of generated text is a difficult problem for large language models.
Approach: They propose a dataset for multilingual, multifaceted summarization evaluation.
Outcome: The proposed dataset can be used to train multilingual summarization systems . it shows that the dataset performs well on the out-of-domain meta-evaluation benchmarks TRUE and mFACE .
MSˆ2: Multi-Document Summarization of Medical Studies (2021.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for multi-document summarization (MDS) are either in the general domain, such as WikiSum, or very small such as DUC 1 or TAC 2011 . Existing systems for summarizing biomedical literature take 1-2 years to complete .
Approach: They propose to use a multi-document summarization system based on BART to assess the quality of the summarized biomedical literature.
Outcome: The proposed system has high summarization quality, but significant work remains to achieve it.
TeSum: Human-Generated Abstractive Summarization Corpus for Telugu (2022.lrec-1)

Copied to clipboard

Challenge: a number of recent datasets for summarisation, scraped the web-content relying on the assumption that summary is made available with the article by the publishers.
Approach: They propose a pipeline that crowd-sources summarization data and then aggressively filters the content via: automatic and partial expert evaluation.
Outcome: The proposed pipeline can be applied to scraped datasets to extract better quality articles-summaries pairs.
Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have enabled them to process increasingly longer sequences, ranging from 2K to 2M tokens and even beyond.
Approach: They propose a synthetic dataset in the financial domain that integrates Chain-of-Thought reasoning into LLMs in a supervised manner to facilitate effective long-context understanding.
Outcome: The proposed model outperforms standard GPT-4o-mini on the Loong benchmark and fine tunes LLaMA-3.1-8B-Instruct on the model, achieving a 28.0% gain on the financial subset.
Alleviating Exposure Bias via Multi-level Contrastive Learning and Deviation Simulation in Abstractive Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Abstractive summarization systems have a severe mismatch between training and inference, i.e., exposure bias.
Approach: They propose a multi-level contrastive learning framework for abstractive summarization and a tailored sparse decoder self-attention pattern to bridge the gap between training and inference.
Outcome: The proposed framework outperforms the state-of-the-art models on two summarization datasets while adding relatively low overhead.
SUM-QE: a BERT-based Summary Quality Estimation Model (D19-1)

Copied to clipboard

Challenge: SUM-QE is a quality estimation model for summarization that captures linguistic qualities that traditional evaluation metrics fail to capture.
Approach: They propose a new quality estimation model based on BERT that addresses linguistic quality aspects that are only indirectly captured by content-based approaches to summary evaluation without comparison with human ratings.
Outcome: The proposed model outperforms existing models on linguistic quality aspects that are only indirectly captured by content-based summarization evaluations without comparison with human ratings.
A Bag of Tricks for Dialogue Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Using a pretrained sequence-to-sequence language model, we explore speaker name substitution, negation scope highlighting, multi-task learning with relevant tasks, and pretraining on in-domain data.
Approach: They propose a pretrained sequence-to-sequence language model that can handle different parts of dialogue belonging to multiple speakers and combine them to produce a coherent monologue summary.
Outcome: The proposed techniques outperform baseline models on a dialogue summarization dataset.
Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to improve output quality without aggregating input tokens are limited by the complexity of aggregation of responses.
Approach: They propose to extract and integrate segment-level commonalities from candidate samples to enhance performance of LLMs in open-ended and reasoning tasks.
Outcome: The proposed method improves performance on reasoning, code generation and mathematical reasoning tasks without requiring additional models and overlooking the knowledge present among the candidates.
What’s in Your Head? Emergent Behaviour in Multi-Task Transformer Models (2021.emnlp-main)

Copied to clipboard

Challenge: Existing paradigms for multi-task training involve a shared pre-trained language model and a small, thin network (head) given an input, a target head is the head that is selected for outputting the final prediction.
Approach: They examine the behaviour of non-target heads when given input that belongs to a different task than the one they were trained for.
Outcome: The non-target heads exhibit emergent behaviour, which may explain the target task, or generalize beyond their original task.
Intrinsic Evaluation of Summarization Datasets (2020.emnlp-main)

Copied to clipboard

Challenge: Almost all popular summarization datasets do not come with inherent quality assurance guarantees.
Approach: They propose to use 5 metrics to evaluate quality of summarization datasets . they find that data usage in recent summarizing research is inconsistent with the properties of the data.
Outcome: The proposed metrics can be inexpensive heuristics for detecting generically low quality examples.
Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors (2023.acl-long)

Copied to clipboard

Challenge: Abstractive summarization systems still include factual errors in generated summaries despite recent improvements in factuality detection .
Approach: They aggregate factuality error annotations from nine existing datasets and stratify them according to the underlying summarization model.
Outcome: The proposed method improves on the ChatGPT-based model and shows that it is not superior for all error types.
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating factual consistency are primarily designed for short summaries of isolated code snippets.
Approach: They propose a reference-free and fine-grained method for evaluating factual consistency in real-world code summaries.
Outcome: The proposed method achieves highest correlation with human judgment among 13 baselines, improving 15-18% over the previous state-of-the-art.
Multi-source Meta Transfer for Low Resource Multiple-Choice Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Existing MCQA datasets are small in size, which increases difficulty of model learning and generalization.
Approach: They propose a multi-source meta transfer framework for low-resource multiple-choice question answering . they extend meta learning by incorporating multiple training sources to learn a generalized feature representation across domains .
Outcome: The proposed framework is independent of backbone language models and can bridge the distribution gap between training sources and target.
From News to Summaries: Building a Hungarian Corpus for Extractive and Abstractive Summarization (2024.lrec-main)

Copied to clipboard

Challenge: Existing models and datasets for training summarization models are limited for less resourceful languages like Hungarian .
Approach: They propose to use a Hungarian corpus for training abstractive and extractive summarization models by cleaning, preprocessing and deduplication.
Outcome: The proposed model trains abstractive and extractive summarization models using the dataset . it will be made publicly available, encouraging replication, further research, and real-world applications across various domains.
Hybrid Alignment Training for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing approaches to align large language models with instructions and preferences are conflicting . et al., 2023b) show that hybrid alignment training can outperform baselines .
Approach: They propose a hybrid alignment training approach based on alternating alignment and modified elastic weight consolidation methods to achieve better collaboration between different alignment tasks.
Outcome: The proposed approach outperforms baseline alignment training methods on summarization and dialogue tasks.
Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination Trends (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have improved summarization, but they still face a challenge of hallucination.
Approach: They propose a taxonomy of errors to address the problem of hallucination in LLMs . they propose two prompt-based approaches for fine-grained error detection .
Outcome: The proposed model outperforms existing metrics in identifying the novel "Contextual Inference" error type.
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs (2025.findings-acl)

Copied to clipboard

Challenge: a core part of legal work that has been underexplored in Legal NLP is the writing and editing of legal briefs.
Approach: They propose to use large language models to help legal professionals with writing briefs by capturing and evaluating their abilities in language models.
Outcome: The proposed tasks show that the models perform well on arguments summarization, argument completion, and case retrieval tasks.
Can Large Language Models Fix Data Annotation Errors? An Empirical Study Using Debatepedia for Query-Focused Text Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: Debatepedia dataset limited by noise and most queries do not have relevance to document .
Approach: They harness the language generation capabilities of two LLMs to regenerate queries in a Debatepedia dataset.
Outcome: The proposed model can regenerate queries from the Debatepedia dataset.
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews (2024.acl-long)

Copied to clipboard

Challenge: Scientific peer review is essential for the quality of academic publications.
Approach: They propose a method that summarises scholarly reviews using a Rational Speech Act framework and novel uniqueness scores.
Outcome: The proposed method generates more discriminative summaries than baseline methods in terms of human evaluation while achieving comparable performance with these methods in term of automatic metrics.
FaithLens: Detecting and Explaining Faithfulness Hallucination (2026.findings-acl)

Copied to clipboard

Challenge: Recent progress in large language models (LLMs) has revolutionized text generation.
Approach: They propose a faithfulness hallucination detection model that can provide binary predictions and corresponding explanations to improve trustworthiness.
Outcome: The proposed model outperforms advanced models on 12 diverse tasks.
Structured Pruning for Efficient Generative Pre-trained Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Large-scale generative Pre-trained Language Models (PLMs) are limited in their deployment in real-world applications.
Approach: They propose to prune the feed-forward networks of generative pre-trained language models to smaller widths without designing extra operators.
Outcome: The proposed method achieves 1.51x/6.96x inference speedup on GPU/CPU with 67% size reduction.
IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Lack of publicly available NLG benchmarks for low-resource languages poses a challenge . authors show that IndoBART and IndoGPT achieve competitive performance on all tasks .
Approach: They propose a benchmark to measure natural language generation progress in three low-resource languages of Indonesia . they use a corpus of pretraining datasets to build their models .
Outcome: The proposed benchmark measures progress in Indonesian, Javanese, and Sundanese . the results highlight the importance of pretraining on closely related, localized languages .
From Sights to Insights: Towards Summarization of Multimodal Clinical Documents (2024.acl-long)

Copied to clipboard

Challenge: a recent WHO report highlights a drastic doctor-to-patient ratio . telehealth is one of the most impactful sectors where AI advances can bring a significant revolution .
Approach: They propose an image-guided encoder-decoder model that uses contextual attention to create detailed visual-guides for multimodal documents.
Outcome: The proposed model outperforms state-of-the-art models on multimodal question and dialogue summarization tasks.
Socratic Pretraining: Question-Driven Pretraining for Controllable Summarization (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to control document controllable summarization lack abundant labeled data.
Approach: They propose a question-driven, unsupervised pretraining objective to improve controllability in document controllable summarization tasks.
Outcome: The proposed method outperforms pre-finetuning approaches on QMSum and SQuALITY.
BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics (2023.acl-long)

Copied to clipboard

Challenge: Existing benchmarks measure the correlation with human judgements of faithfulness on model-generated summaries, but they are insufficient for diagnosing whether metrics are consistent, effective on human-written texts, and sensitive to different error types.
Approach: They propose to use unfaithful minimal pairs to measure the consistency of automatic faithfulness metrics by comparing human-written summary pairs with a dataset of 889 human-writing, minimally different summary pairs.
Outcome: The proposed benchmarks show that the most discriminative metrics tend not to be the most consistent, and that the best performing metrics are sensitive to errors.
Prediction-Augmented Generation for Automatic Diagnosis Tasks (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) adopt autoregressive architecture, predicting the next word token based on the preceding context.
Approach: They propose a method that integrates task-specific predictive models as external tools to improve model generation quality and accuracy.
Outcome: The proposed method improves the generation quality and predictive accuracy of large language models in inference-driven tasks.
Context Filtering with Reward Modeling in Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Question Answering (QA) tasks require a mix of relevant and irrelevant information in these contexts to perform well.
Approach: They propose a context filtering approach that removes non-essential details, summarizing crucial content through Reward Modeling.
Outcome: The proposed approach outperforms baseline models in 6.8-folds.
ManiTweet: A New Benchmark for Identifying Manipulation of News on Social Media (2025.coling-main)

Copied to clipboard

Challenge: Existing studies have focused on the identification of social media posts that contain misrepresentations of information within associated news articles.
Approach: They propose a data collection schema and curated a dataset called ManiTweet, consisting of 3.6K pairs of tweets and corresponding articles.
Outcome: The proposed model outperforms large language models on the ManiTweet dataset and reveals intriguing connections between manipulation and the domain and factuality of news articles.
BARThez: a Skilled Pretrained French Sequence-to-Sequence Model (2021.emnlp-main)

Copied to clipboard

Challenge: Inductive transfer learning has taken the entire NLU field by storm, with models such as BERT and BART setting new state-of-the-art on countless tasks.
Approach: They introduce a large-scale pretrained seq2seq model for French that is very competitive with state-of-the-art BERT-based French language models such as CamemBERT and FlauBERT.
Outcome: The proposed model outperforms existing models on discriminative and generative tasks on a French summarization dataset.
ARMAN: Pre-training with Semantically Selecting and Reordering of Sentences for Persian Abstractive Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Abstractive summarization is one of the areas influenced by pre-trained language models.
Approach: They propose a Transformer-based encoder-decoder model pre-trained with three novel objectives to address this issue.
Outcome: The proposed model outperforms previous models on six Persian summarization tasks . it also outperformed previous models in textual entailment, question paraphrasing, and question answering .
Narrative License and Model Sycophancy in LLM Summaries of Scientific Work (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used to summarize academic work . however, they can exaggerate or mischaracterize findings .
Approach: They examine how Narrative License (NL) emerges in large language models summaries . authors find that stated stances and user personas produce predictable shifts .
Outcome: The proposed models can exaggerate or mischaracterize findings in scholarly articles . the authors show that the models' "sycophancy" can reduce NL .
Re-evaluating Evaluation in Text Summarization (2020.emnlp-main)

Copied to clipboard

Challenge: Automated evaluation metrics are an essential part of the development of text-generation tasks such as summarization.
Approach: They propose to use top-scoring system outputs to assess the reliability of automatic evaluation metrics for text summarization.
Outcome: The proposed evaluation method is based on human judgments from 25 top-scoring neural summarization systems.
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable capabilities in tasks such as summarization, arithmetic reasoning, and question answering.
Approach: They propose a framework that explores decisions’ consequences from multiple stakeholder perspectives and a SKIG framework to enhance moral reasoning in large language models.
Outcome: The proposed framework exhibits marked improvements compared to baselines across different language models and benchmarks.
MISP-Meeting: A Real-World Dataset with Multimodal Cues for Long-form Meeting Transcription and Summarization (2025.acl-long)

Copied to clipboard

Challenge: Existing systems that can recognize spoken content, extract key information, and produce concise summaries are lacking in meeting transcription and summarization.
Approach: They propose a multimodal dataset that integrates information from speech, vision, and text modalities to facilitate automatic meeting transcription and summarization (AMTS).
Outcome: The proposed dataset reduces the character error rate (CER) by 36.60% to 20.27% and improves speech recognition and large language models.
Pre-training Language Models for Comparative Reasoning (2023.emnlp-main)

Copied to clipboard

Challenge: Comparative reasoning is a process of comparing objects, concepts, or entities to draw conclusions.
Approach: They propose a framework to pre-train language models for enhancing comparative reasoning abilities . they collect scalable data for text-based entity comparison .
Outcome: The proposed framework significantly improves comparative reasoning abilities under low-resource conditions on downstream tasks.
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency (2024.lrec-main)

Copied to clipboard

Challenge: Abstractive summarization models (LLMs) have demonstrated impressive performance in various tasks, but they are still suffering from factual inconsistency problem called hallucination.
Approach: They propose to improve the faithfulness of large language models by impelling them to process the entire article more fairly and faithfully.
Outcome: The proposed strategy improves the faithfulness of large language models in summarization while maintaining their fluency and informativeness.
Multilingual Large Language Models Are Not (Yet) Code-Switchers (2023.emnlp-main)

Copied to clipboard

Challenge: Existing multilingual Large Language Models are not specifically trained with objectives for managing code-switching scenarios.
Approach: They propose to use multilingual Large Language Models to perform sentiment analysis, machine translation, summarization and word-level language identification to compare their performance to fine-tuned models of much smaller scales.
Outcome: The proposed models show that they underperform in comparison to fine-tuned models of much smaller scales.
Evaluating Factuality in Cross-lingual Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for monolingual summarization require translation to evaluate the factuality of cross-lingual summmarization.
Approach: They propose to analyze cross-lingual factuality by collecting annotations and generated summaries from models at summary level and sentence level.
Outcome: The proposed dataset shows that over 50% of generated summaries contain factual errors with different characteristics from monolingual summarization.
Towards Understanding Omission in Dialogue Summarization (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for dialogue summarization are far from satisfactory . omission is a major factor in affecting the quality of summarizing, but few studies have explored the problem .
Approach: They propose a dataset that provides high-quality omission labels for dialogue summarization . they propose to use this dataset to detect omitted dialogue utterances .
Outcome: The proposed dataset improves summarization quality by providing ground-truth omission labels . the proposed dataset and codes are publicly available .
Uniform Complexity for Text Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models do not capture factors that contribute to producing consistent text.
Approach: They propose a benchmark test to evaluate text complexity in generative models by observing linguistic properties of input prompts.
Outcome: The proposed model fails to preserve complexity of input prompts even if finetuned with professionally written texts.
Revisiting the Architectures like Pointer Networks to Efficiently Improve the Next Word Distribution, Summarization Factuality, and Beyond (2023.findings-acl)

Copied to clipboard

Challenge: Existing solutions for word probability distributions are limited and the output softmax layer is inherently limited.
Approach: They propose to use the output softmax layer to compute the word probability distribution instead of using pointer networks to break the bottleneck.
Outcome: The proposed method improves factCC score by 2 points in CNN/DM and XSUM dataset, and MAUVE scores by 30% in bookSum paragraph-level dataset.
Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting Transcripts (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to meeting summarization are limited due to noise, lengthy transcripts, and scattered salient information.
Approach: They propose a two-step framework for meeting summarization that leverages a self-supervised paradigm to reconstruct transcripts and a relative positional bucketing algorithm to equip models to generate the summary.
Outcome: The proposed method significantly reduces memory consumption and processing time on two meeting summarization datasets.
Improving Faithfulness by Augmenting Negative Summaries from Fake Documents (2022.emnlp-main)

Copied to clipboard

Challenge: Current abstractive summarization systems tend to hallucinate unfaithful content . however, the most common method does not disentangle factual errors from other errors.
Approach: They propose a back-translation-style approach to augment negative samples that mimic factual errors made by the model.
Outcome: The proposed method improves faithfulness without sacrificing informativeness . it incorporates negative samples into training, and produces faithful/unfaithful summaries .
Align then Summarize: Automatic Alignment Methods for Summarization Corpus Creation (2020.lrec-1)

Copied to clipboard

Challenge: Summarizing text is not a straightforward task.
Approach: They propose to use automated transcriptions to generate reports from automatic transcriptions as a dataset for neural summarization.
Outcome: The proposed model improves on publicmeetings corpus on a dataset of aligned public meetings.
MeetingQA: Extractive Question-Answering on Meeting Transcripts (2023.acl-long)

Copied to clipboard

Challenge: Meeting transcripts are a promising domain for natural language tasks . lack of annotated data impedes research on other important tasks in this domain .
Approach: They propose an extractive QA dataset comprising questions asked by meeting participants and corresponding responses.
Outcome: The proposed dataset extracts questions asked by meeting participants and corresponding responses from transcripts.
Learning to Rank Salient Content for Query-focused Summarization (2024.emnlp-main)

Copied to clipboard

Challenge: Query-focused summarization (QFS) is gaining prominence in research community.
Approach: They propose to integrate Learning-to-Rank (LTR) with Query-focused Summarization (QFS) to enhance the summary relevance via content prioritization.
Outcome: The proposed model outperforms the state-of-the-art on QMSum benchmark and SQuALITY benchmark while offering a lower training overhead.
Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree Search (2026.acl-long)

Copied to clipboard

Challenge: Controllable summarization is a form of outputs that tailors summaries to user-specified attributes.
Approach: They propose an adaptive planning framework that reframes the task as planning the order of sequential attribute control with a customized Monte Carlo Tree Search.
Outcome: The proposed framework surpasses LLM-based self-planning models and fine-tuned baselines in multi-attribute controllable summarization.
Identifying Factual Inconsistencies in Summaries: Grounding LLM Inference via Task Taxonomy (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have focused on specialized BERT-variants and recent LLMs to reason inconsistencies.
Approach: They propose to incorporate task-specific taxonomy into inferences to facilitate both zero-shot and supervised paradigms.
Outcome: The proposed model outperforms specialized non-LLM and recent LLM models in a number of domains.
SIMBA UQ: Similarity-Based Aggregation for Uncertainty Quantification in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Uncertainty quantification (UQ) provides measures of uncertainty, such as an estimate of the confidence in an LLM’s generated output.
Approach: They propose a black-box approach where consistency is used as a proxy for confidence in a model's output.
Outcome: The proposed methods are primarily but not necessarily entirely black- box, with consistency between output and other sampled generations used as a proxy for confidence in its correctness.
RISE: Leveraging Retrieval Techniques for Summarization Evaluation (2023.findings-acl)

Copied to clipboard

Challenge: Summarization evaluation approaches have relied on ROUGE for summarization, but they fall short of human evaluations.
Approach: They propose a new approach to evaluate summaries by leveraging retrieval techniques . they use a dual-encoder retrieval setup to train a retrieval task .
Outcome: The proposed method outperforms existing methods on two document summarization benchmarks and a long document summmarization test.
CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code (2023.emnlp-main)

Copied to clipboard

Challenge: Current work on understanding assembly code is oriented towards generating function names, which involve numerous abbreviations that make them confusing.
Approach: They propose a control flow graph and pseudo code guided binary code summarization framework to learn the comprehensive binary function execution behavior and logic semantics.
Outcome: The proposed framework improves the efficiency of reverse engineering on 3 different binary optimization levels for 3 different computer architectures.
Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding (2025.emnlp-main)

Copied to clipboard

Challenge: In-Context Learning (ICL) is a key method in prompt engineering, but its long retrieved contexts and limited token throughput will slow reasoning speeds.
Approach: They propose a method that leverages the overlap between context and model output to generate drafts from the context.
Outcome: The proposed method achieves the highest mean speedup on Vicuna-7B, Llama2-7B-Chat, and Llma3-8B-Instruct tasks.
Unveiling the Essence of Poetry: Introducing a Comprehensive Dataset and Benchmark for Poem Summarization (2023.emnlp-main)

Copied to clipboard

Challenge: Summarization of poetry is a challenging task as it can be easily lost if only the literal meaning is considered.
Approach: They propose to use poetry as a model to summarize poetry and provide a dataset to evaluate their creative language interpretation capacity.
Outcome: The proposed dataset consisting of 3011 samples and its corresponding summarized interpretation in the English language provides an opportunity to evaluate the creative language interpretation capacity of the proposed models.
Small LLMs Are Weak Tool Learners: A Multi-LLM Agent (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized natural language processing with impressive capabilities, but they lack domain specificity, real-time information and face challenges in solving specialized problems.
Approach: They propose a multi-LLM approach that decomposes the aforementioned capabilities into a planner, caller, and summarizer.
Outcome: The proposed model outperforms existing models by demonstrating its effectiveness and advantages in tool learning.
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization (2025.acl-long)

Copied to clipboard

Challenge: n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear.
Approach: They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics.
Outcome: The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand.
Adaptive LoRA Merge with Parameter Pruning for Low-Resource Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for adapting LLMs to low-resource tasks keep LoRA parameters frozen and the low-level problem out of their scope.
Approach: They propose a LoRA merge method that updates and prunes LoRA parameters through fine-tuning with minimal target task data.
Outcome: The proposed method improves performance on a low-resource language generation task and improves on previous methods.
Assessing Privacy Risks in Language Models: A Case Study on Summarization Tasks (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models have revolutionized the field of NLP by achieving state-of-the-art performance on various tasks.
Approach: They investigate the membership inference attack by using model's API to determine if a sample was part of the training data.
Outcome: The proposed model is able to identify if a sample was part of the training data and exploits its similarity and resistance to document modifications as potential MI signals on widely used datasets.
Multilingual Generation in Abstractive Summarization: A Comparative Study (2024.lrec-main)

Copied to clipboard

Challenge: Existing models for multilingual generation lack thorough analysis due to extensive linguistic diversity.
Approach: They propose to classify multilingual generation methodologies into three categories based on their underlying modeling principles . they introduce an automatic metric to mitigate spurious correlations associated with language mixing .
Outcome: The proposed model improves in high-resource, low-resourced, and zero-shot scenarios.
Multi-Objective Forward Reasoning and Multi-Reward Backward Refinement for Product Review Summarization (2024.lrec-main)

Copied to clipboard

Challenge: Product review summarization aims to generate a concise summary based on product reviews . factual accuracy, aspect comprehensiveness, and content relevance are challenges .
Approach: They propose an FB-Thinker framework to improve product review summarization ability . they propose two Chinese product review summary datasets for instruction-tuning and evaluation .
Outcome: The proposed framework improves product review summarization with forward reasoning and backward refinement.
NormAL LoRA: What is the perfect size? (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are crucial for enabling intelligent experiences across applications.
Approach: They propose a low-rank adaptive localization method that uses rank-norm regularization to determine the optimal rank for each weight matrix.
Outcome: NormAL LoRA reduces adapter parameters by 37% while preserving full fine-tuning performance.
Re-Evaluating Evaluation for Multilingual Summarization (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that automated evaluation approaches correlate with human ratings in English, but this is unclear for other languages.
Approach: They construct a small-scale pilot dataset containing article-summary pairs and human ratings in English, Chinese and Indonesian to measure the strength of summaries.
Outcome: The results show that standard metrics are unreliable measures of quality in Chinese and Indonesian.
Towards Aligning Language Models with Textual Feedback (2024.emnlp-main)

Copied to clipboard

Challenge: Using textual feedback, language models can be trained to learn from textual inputs.
Approach: They propose an approach that aligns language models with user preferences expressed in text.
Outcome: The proposed approach outperforms PPO on toxicity reduction, summarization, and dialog response tasks while achieving the same performance with only 20% of the samples.
PSentScore: Evaluating Sentiment Polarity in Dialogue Summarization (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have focused on summarizing factual information, leaving out affective content.
Approach: They propose to quantify the preservation of affective content in dialogue summaries using PSentScore measures.
Outcome: The proposed measures show that state-of-the-art summarization models do not preserve well affective content in their summaries.
Quantifying the Impact of Disfluency on Spoken Content Summarization (2024.lrec-main)

Copied to clipboard

Challenge: a recent study has found that disfluencies negatively impact spoken content summarization .
Approach: They aim to quantify the impact of disfluency on spoken content summarization . they also investigate two methods towards improving summarizing in the presence of disflouencies .
Outcome: The proposed methods improve summarization quality in the presence of disfluencies.
COVER: Context-Driven Over-Refusal Verification in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have become increasingly prevalent in the field of Natural Language Processing (NLP), achieving unprecedented performance across linguistic tasks.
Approach: They propose a framework to quantify and analyze context-driven over-refusal . they find that over-fusals depend on the task, system prompts, model family, and the number of retrieved documents.
Outcome: The proposed framework quantifyes and analyzes the concept of context-driven over-refusal on two public corpora.
Exploration-Driven Reinforcement Learning for Expert Routing Improvement in Mixture-of-Experts Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: MoE-based LLMs are not explicitly supervised to select suitable experts.
Approach: They propose Exploration-Driven Reinforcement Learning (ERL) which explicitly optimizes the router by exploration of alternative routing paths.
Outcome: The proposed method improves summarization (SAMSum, XSUM, question answering, and language modeling), and raises routing quality, delivering 8.9 higher MRR than baselines over 100 perturbed routing paths.
Multi-document Summarization through Multi-document Event Relation Graph Reasoning in LLMs: a case study in Framing Bias Mitigation (2025.acl-long)

Copied to clipboard

Challenge: a recent study has focused on detecting media bias in news articles . a multi-document event relation graph is used to generate a neutralized summary .
Approach: They propose to generate a neutralized summary given multiple articles presenting different ideological views.
Outcome: The proposed method mitigates media bias and improves content preservation.
Read As Human: Compressing Context via Parallelizable Close Reading and Skimming (2026.acl-long)

Copied to clipboard

Challenge: Existing task-aware methods require loading the entire input sequence at once for compression, which suffer from computational inefficiency.
Approach: They propose a framework that adopts an adaptive hybrid reading strategy to reduce computational inefficiency and redundant information in long-context scenarios.
Outcome: Experiments show that RAM outperforms baselines on multiple question answering and summarization benchmarks while delivering up to a 12x speedup on long inputs.
GMSA: Enhancing Context Compression via Group Merging and Layer Semantic Alignment (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved remarkable performance across NLP tasks . however, in long-context scenarios, they face high computational cost and information redundancy.
Approach: They propose an encoder-decoder context compression framework that generates a compact sequence of soft tokens for downstream tasks.
Outcome: Experiments show that GMSA outperforms baselines on multiple long-context question answering and summarization benchmarks while maintaining low end-to-end latency.
Summarizing Speech: A Comprehensive Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Podcasts and other audiovisual content are becoming more and more a part of everyday communication and the digital age is changing from text to voice.
Approach: They synthesize the current state of the field and highlight the need for realistic evaluation benchmarks and multilingual datasets.
Outcome: The proposed frameworks are based on evaluation protocols and datasets and highlight the need for realistic benchmarks and multilingual datasets.
Principled Content Selection to Generate Diverse and Personalized Multi-Document Summaries (2025.acl-long)

Copied to clipboard

Challenge: Large language models exhibit the _”lost in the middle” phenomenon when they are unevenly attending to different parts of the provided context.
Approach: They propose principled content selection as a way to increase source coverage . they use determinantal point processes to prioritize diverse content .
Outcome: The proposed method improves source coverage on the DiverseSumm benchmark.
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination (2025.emnlp-main)

Copied to clipboard

Challenge: despite LLMs becoming increasingly multilingual, most studies on detecting and quantifying LLM hallucination are English-centric .
Approach: They train a multilingual hallucination detection model and conduct a large-scale study across 30 languages and 6 open-source LLM families.
Outcome: The proposed model is based on an English-centric model and annotates gold data for five high-resource languages.
How Private are Language Models in Abstractive Summarization? (2025.emnlp-main)

Copied to clipboard

Challenge: Effective protection of private information is essential for knowledge dissemination in sensitive domains such as medical and legal.
Approach: They perform a comprehensive study of privacy risks in LM-based summarization using closed- and four-weight models of different sizes and families.
Outcome: The proposed models show that they leak personally identifiable information in their summaries, compared to human-generated summary summators, which show significantly higher privacy protection levels.
Attn-GS: Attention-Guided Context Compression for Efficient Personalized LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to personalize large language models (LLMs) rely on heuristic methods to compress user profiles but they ignore how LLMs process and prioritize different profile components.
Approach: They propose an attention-guided context compression framework that leverages attention feedback from a marking model to mark important personalization sentences and guides a compression model to generate task-relevant compressed user contexts.
Outcome: The proposed framework outperforms baselines across tasks, token limits, and settings while reducing token usage by 50 times.
Making Revisions Understandable: A Survey of Edit Intentions, Methods, and Applications (2026.findings-acl)

Copied to clipboard

Challenge: Text revision is a core process in document creation, capturing how authors iteratively refine, reorganize, and improve written content.
Approach: They synthesize text revision research through the lens of edit intentions . they review prior work across the revision workflow including corpus construction, edit intention taxonomies, edit intentions, and edit intention identification.
Outcome: The proposed approach synthesizes datasets, taxonomies, identification methods, and applications and highlights key open research directions.
Calibrating Model-Based Evaluation Metrics for Summarization (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in summary evaluation are based on model-based metrics to assess quality dimensions, such as completeness, conciseness, and faithfulness.
Approach: They propose a general framework that generates individual and average proxy scores without relying on reference summaries, human annotations, or expensive model-based metrics.
Outcome: The proposed framework outperforms baselines on seven datasets on continuous-value scenarios, such as summarization, but is applicable to discrete-value tasks, such QA.
Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation (2026.acl-long)

Copied to clipboard

Challenge: Traditional metrics like BLEU and BERTScore fail to capture semantic fidelity in generative text-to-text tasks.
Approach: They propose a cross-examination framework that generates verifiable questions from each text and performs a Cross-exam to derive three interpretable scores: Coverage, Conformity, and Consistency.
Outcome: The proposed framework detects critical errors across translation, summarization and clinical note-generation and human expert validation shows it is reliable without gold references.
Learning to Control Summaries with Score Ranking (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in summarization focus on improving summary quality across multiple dimensions, but they overlook the challenge of controlling summary generation with respect to individual dimensions.
Approach: They propose a loss function that aligns model outputs with fine-grained, model-based evaluation scores to enable both improvement in summary quality and dimension-specific control.
Outcome: The proposed method improves the overall quality of summaries while maintaining strong control over individual quality dimensions.
Benchmarking Agentic Newswriting via Journalistic Workflows (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in autonomous digital agents highlight their potential for structured tasks through autonomous decision-making and task decomposition, but it remains unclear how well such systems support real-world information-intensive workflows.
Approach: They propose a benchmark to evaluate how journalists can use agents to organize and organize information from the web.
Outcome: The proposed system can be used to iterate and evaluate newswriting tasks in real-world situations.
Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization (2026.acl-long)

Copied to clipboard

Challenge: a new approach to adapt generalist models to expert domains is needed to overcome this problem.
Approach: They propose a parameter-efficient domain adaptation approach that combines vocabulary adaptation with pretraining for LLM-based text summarization.
Outcome: The proposed approach reduces training time by 35-55% over continual pretraining and reduces parameter counts up to 37% w.r.t expansion-only methods.
PROBE: PROcess-Based BEnchmark for Hallucination Detection (2026.findings-acl)

Copied to clipboard

Challenge: Existing agentic applications rely on LLMs to self-assess the factuality of outputs . but current LLM systems fail to detect hallucinations .
Approach: They propose a benchmark that breaks down hallucination detection into four critical steps . they show that when halluciation detection is treated as a multi-step process, all models achieve considerably better performance.
Outcome: The proposed benchmark breaks down hallucination detection into four critical steps . it shows that when halluciation detection is treated as a multi-step process, all models achieve considerably better performance.
Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding (2026.findings-acl)

Copied to clipboard

Challenge: Large language models produce content that contradicts or overlooks information provided in the input context, a phenomenon known as faithfulness hallucination.
Approach: They propose a lightweight framework that boosts the generation probability of context-relevant tokens by boosting the generation of tokens.
Outcome: The proposed framework improves faithfulness metrics with minimal generation overhead.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations