Papers with summarization
Copied to clipboard
| Challenge: | Position bias is a tendency of a model unfairly prioritizing information from certain parts of the input text over others, leading to undesirable behavior. |
| Approach: | They propose to measure position bias in large language models for zero-shot summarization tasks by measuring position bias. |
| Outcome: | The proposed model performance and position biases lead to new insights and discussion on zero-shot summarization tasks. |
Copied to clipboard
| Challenge: | Existing summarization systems struggle to address diverse linguistic and cognitive barriers among general readers. |
| Approach: | They propose a multi-agent framework that integrates template-based planning with an iterative feedback loop guided by simulated readers and domain expert revision to address comprehension barriers such as unknown terms, missing contexts, and confusing sentences. |
| Outcome: | The proposed framework improves readability and factuality across multiple datasets and human evaluations show that it is more accessible to a wide range of readers. |
Copied to clipboard
| Challenge: | Existing evaluation strategies for analyzing economic data with narratives are limited due to the complexity of the interplay of numerous factors and the difficulty in isolating causal relationships. |
| Approach: | They propose to use two Twitter datasets to capture economy-related narratives and use them to construct models using large language models. |
| Outcome: | The proposed models are able to predict macroeconomic fluctuations using the extracted or extracted narratives in two Twitter datasets. |
Copied to clipboard
| Challenge: | Timeline EXtraction library provides Java implementation of TimeML annotations and tools for programmatic manipulation of Timeline graphs. |
| Approach: | jTLEX provides a Java implementation of TimeLine EXtraction algorithm and utilities for programmatic manipulation of TimeML graphs. |
| Outcome: | jTLEX provides a Java implementation of the TimeLine EXtraction algorithm, along with utilities for programmatic manipulation of TimeML graphs. |
Copied to clipboard
| Challenge: | Using a novel task, we advocate automatic pull quote selection to engage readers with thought-provoking articles . pull quotes increase enjoyment and readability, shape reader perceptions, and facilitate learning. |
| Approach: | They propose a task that automatically selects pull quotes from articles with more salient presentation. |
| Outcome: | The proposed task differs from similar tasks such as summarization and clickbait identification by several aspects. |
Copied to clipboard
| Challenge: | Summarizing sales calls is a routine task performed manually by salespeople. |
| Approach: | They propose a production system which combines generative models fine-tuned for customer-agent setting, with a human-in-the-loop user experience for an interactive summary curation process. |
| Outcome: | The proposed system can handle training data scarcity and privacy constraints in an industrial setting. |
Copied to clipboard
| Challenge: | Empirical results show that a modified beam decoding implementation improves decoding performance of strong, neural language generation models. |
| Approach: | They propose a modification to a beam decoding implementation that generalizes the stopping criterion and provides flexibility to the depth of search. |
| Outcome: | The proposed method improves decoding performance of strong models on news text summarization and machine translation over diverse language pairs with negligible inference slowdown. |
Copied to clipboard
| Challenge: | Student Evaluations of Teaching (SETs) are used in colleges and universities to assess student perceptions about their courses. |
| Approach: | They propose a system that leverages sentiment analysis, aspect extraction, summarization and visualization techniques to provide organized illustrations of SET findings to instructors and other reviewers. |
| Outcome: | The proposed system can be used by 10 professors from diverse departments to analyze SET results. |
Copied to clipboard
| Challenge: | Recent progress of abstractive text summarization relies on large pre-trained sequence-to-sequence Transformer models, which are computationally expensive. |
| Approach: | They propose to distill large Transformer summarization models into smaller ones with minimal performance loss by manipulating attention temperatures in Transformers. |
| Outcome: | The proposed method outperforms vanilla pseudo-labeling based methods on three summarization datasets and is shorter and more abstractive. |
Copied to clipboard
| Challenge: | Previous work has used task-agnostic pretraining methods like masked language models or corrupted span prediction to improve performance on downstream tasks. |
| Approach: | They propose to use a task-agnostic pretraining to improve on low-resource tasks. |
| Outcome: | The proposed model can predict extracted gap sentences on summarization with a low resource and zero shot setup. |
Copied to clipboard
| Challenge: | Existing aspects-based summarization models are domain-specific due to large differences in the type of aspects for different domains. |
| Approach: | They propose a large-scale dataset for multi-domain aspect-based summarization using Wikipedia articles from 20 different domains. |
| Outcome: | The proposed model is based on Wikipedia articles from 20 different domains and uses the section titles and boundaries of each article as a proxy for aspect annotation. |
Copied to clipboard
| Challenge: | Existing studies focus on summarizing news documents or structured documents. |
| Approach: | They propose to use a large-scale narrative summarization dataset to encourage research . they find there is a performance gap between humans and the models on NarraSum . |
| Outcome: | The proposed dataset shows that humans and state-of-the-art models perform poorly when summarizing a narrative . it contains 122K narratives collected from synopses of movies and TV episodes with diverse genres . |
Copied to clipboard
| Challenge: | Large language models (LLMs) can generate fluent text, but the quality of generated content depends on its consistency with the given input. |
| Approach: | They constructed a Japanese evaluation dataset for hallucination detection in summarization by manually annotating sentence-level faithfulness labels in LLM-generated summaries of Japanese documents. |
| Outcome: | The proposed model can detect hallucinations in Japanese documents by annotating faithfulness labels in Japanese summaries. |
Copied to clipboard
| Challenge: | Video Large Language Models (VLLMs) exhibit impressive zero-shot capabilities in video analysis, but their performance varies significantly depending on the LLM prompt, the characteristics of the video, and the properties of the training data and LLM architecture. |
| Approach: | They propose to use Chain-of-Thought prompting to inject knowledge extracted by external, lightweight models into video summarization benchmarks to evaluate their performance. |
| Outcome: | The proposed solutions improve summarization performance by injecting knowledge extracted by external, lightweight models. |
Copied to clipboard
| Challenge: | Existing models for natural language and programming languages are lagging behind due to a lack of large datasets and benchmarks. |
| Approach: | They present a large parallel dataset of Java methods and natural language descriptions that is used to train deep neural models. |
| Outcome: | The proposed dataset improves code summarization and code search by 22% and opens up possibilities for pretrained language models for Java. |
Copied to clipboard
| Challenge: | Existing methods for text augmentation perform data augmentation and downstream tasks separately. |
| Approach: | They propose a framework to perform text augmentation and the downstream task end-to-end. |
| Outcome: | The proposed framework performs text augmentation and the downstream task end-to-end on a text classification dataset. |
Copied to clipboard
| Challenge: | Existing studies on redundancy are focused on salience alone. |
| Approach: | They propose to combine salience and novelty to score redundancy in extractive summarization systems . they also propose to balance saliance and redundancies by scoring redundants first . |
| Outcome: | Empirical results show that AREDSUM-CTX scores salience first, then learns to balance saliency and redundancy. |
Copied to clipboard
| Challenge: | Scientific publications are becoming more multimedia, containing both text and visual content. |
| Approach: | They propose a framework for Scientific Multimodal Summarization with Multimodal Output . it leverages the power of large language models and extends its capability to cross-modal understanding . |
| Outcome: | The proposed framework outperforms uni- and multi-modality methods on two new datasets . it leverages the power of large language models and extends its capability to cross-modal understanding . |
Copied to clipboard
| Challenge: | Automatic text summarization is the task of generating a summary of a long text by condensing it to its most important parts. |
| Approach: | They propose a tool to visually explore document summarization systems based on three well-known summary quality criteria . |
| Outcome: | The proposed tool compiles outputs of 55 state-of-the-art document summarization approaches and visually explores them during a qualitative assessment. |
Copied to clipboard
| Challenge: | Current approaches to legal summarization struggle with content theme deviation and inconsistent writing styles due to the content of the source document. |
| Approach: | They propose a retrieval-augmented framework that utilizes exemplar summaries along with the source document to guide the model. |
| Outcome: | The proposed model outperforms models that do not utilize exemplars and those that rely on similarity-based exemplar selection. |
Copied to clipboard
| Challenge: | Texar is an open-source text generation toolkit that supports a broad set of text generation tasks. |
| Approach: | They introduce Texar, an open-source text generation toolkit that supports text generation tasks. |
| Outcome: | Texar supports machine translation, summarization, dialog, content manipulation, and more. |
Copied to clipboard
| Challenge: | Existing approaches to generative AI for large language models struggle when executing complex tasks simultaneously. |
| Approach: | They propose a novel approach tailored specifically for compositional multi-tasking scenarios . they add a learnable projection layer on top of the combined summarization and translation adapters. |
| Outcome: | The proposed approach performs well and is fast in both cloud-based and on-device implementations. |
Copied to clipboard
| Challenge: | Taking excerpts of text can be problematic, as key pieces may not be explicit in a local window. |
| Approach: | They define a problem of sentence decontextualization by rewriting a sentence to be interpretable out of context while preserving its meaning. |
| Outcome: | The proposed method can be used in question answering and document understanding tasks. |
Copied to clipboard
| Challenge: | Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems . |
| Approach: | They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language. |
| Outcome: | The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature. |
Copied to clipboard
| Challenge: | Existing pipelines for generative tasks require extensive manual effort and domain expertise to achieve task-optimal performance. |
| Approach: | They propose a framework bridging discrete and continuous prompt optimization through feedback-guided gradient descent in embedding space. |
| Outcome: | The proposed framework bridges discrete and continuous prompt optimization through feedback-guided gradient descent in embedding space. |
Copied to clipboard
| Challenge: | Existing approaches to text generation combine task descriptions and examples with supervised learning. |
| Approach: | They propose a method for text generation that is based on pattern-exploiting training. |
| Outcome: | The proposed approach improves on several summarization and headline generation datasets. |
Copied to clipboard
| Challenge: | a study examines how to build meeting summarization systems using large language models . closed-source models are generally better in terms of performance, but open-source ones are more advantageous for industrial use . |
| Approach: | They compare closed-source and open-source meeting summarization models for real-world use . they find that closed-sourced models are generally better in terms of performance . however, smaller open-sourced LLMs could still achieve comparable performance if they are open . |
| Outcome: | The proposed model is more efficient for industrial use than closed-source models due to privacy concerns and high cost. |
Copied to clipboard
| Challenge: | Large language models struggle with precise length control, particularly in zero-shot settings. |
| Approach: | They propose to use length approximation, target adjustment, sample filtering and automated revisions to improve LLMs' length control capabilities. |
| Outcome: | The proposed methods improve length control in large language models while maintaining or enhancing summary quality without the need for model fine-tuning or architectural changes. |
Copied to clipboard
| Challenge: | Automated summarization methods are efficient but can suffer from low quality. |
| Approach: | They conducted an experiment with 72 participants to compare post-editing provided summaries with manual summarization for summary quality, human efficiency, and user experience. |
| Outcome: | The results show that post-editing improves summary quality, human efficiency, and user experience on formal (XSum news) and informal (Reddit posts) text. |
Copied to clipboard
| Challenge: | Existing benchmarks for query-focused summarization are small for training large neural models. |
| Approach: | They propose a unified modeling framework for query-focused summarization . they model queries as discrete latent variables over document tokens . |
| Outcome: | The proposed framework outperforms strong comparison systems across benchmarks, query types, document settings, and target domains. |
Copied to clipboard
| Challenge: | Medical text generation systems are widely used to assist with administrative work and highlight salient information to support decision-making. |
| Approach: | They propose a set of metrics to evaluate completeness, conciseness, and attribution of medical text at a fine-grained level. |
| Outcome: | The proposed framework exhibits substantially higher agreement with medical experts than existing metrics. |
Copied to clipboard
| Challenge: | Entity-centric summarization is a type of controllable summarizing that aims to produce a summary specific to a given target entity. |
| Approach: | They propose to recast a sentence selection task as a controllable summarization using a dataset supported by EntSUM. |
| Outcome: | The proposed framework outperforms the current state-of-the-art in the sentence selection task and outperformed the competitive entity-centric Lead 3 heuristic by 1.1 F1. |
Copied to clipboard
| Challenge: | aggregators consume millions of articles every day, making it difficult to quickly identify key events and miss less-reported stories. |
| Approach: | a new kind of summarization engine was needed to condense large volumes of news into short, easy to absorb points. |
| Outcome: | NSTM can be used to summarize news articles in seconds and quickly and efficiently. |
Copied to clipboard
| Challenge: | Summarization of multi-party dialogues is a critical capability in industry . but generating high-quality summaries in practice is challenging . prior work has focused on static datasets and benchmarks, a condition rare in practical scenarios . |
| Approach: | They present an agentic system to summarize multi-party interactions using static datasets. |
| Outcome: | The proposed system can summarize multi-party interactions using a set of complex requirements. |
Copied to clipboard
| Challenge: | Existing methods for generating summarizations using QA-based supervision produce higher quality summaries than baseline methods. |
| Approach: | They propose a method for incorporating question-answering signals into a summarization model by automatically marking document NPs as salient based on whether they are answered in the gold summaries. |
| Outcome: | The proposed method generates higher-quality summaries than baseline methods on benchmark summarization datasets. |
Copied to clipboard
| Challenge: | Abstractive and extractive methods are used to condense long text into concise summaries while retaining essential information. |
| Approach: | They propose to use paper structure to extract paper summaries from long text . they provide a large-scale dataset of COVID-19-related papers . |
| Outcome: | The proposed framework generates more comprehensive and valuable summaries compared to previous work on COVID-19-related papers. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are trained for factual accuracy, but can conflict with the critical demand for source fidelity. |
| Approach: | They propose a reproducible framework to elicit and measure HFH using controlled entity-level perturbations and strategic entity selection. |
| Outcome: | The proposed framework reduces HFH rates by 50% across summarization, rephrasing, and QA tasks. |
Copied to clipboard
| Challenge: | Key Point Analysis (KPA) extracts the main points from opinions and quantifies their prevalence. |
| Approach: | They propose a key point analysis framework that extracts the main points from opinions and quantifies their prevalence. |
| Outcome: | The proposed system is able to match sentences to key points over five datasets and demonstrate its performance. |
Copied to clipboard
| Challenge: | Recent efforts focus on automatic summarization of individual cases, which condense the content of a single case, making it easier for legal professionals to grasp key points. |
| Approach: | They propose a pipeline to generate multi-case structured reports using entire body of case law on user-specified topics within the European Court of Human Rights. |
| Outcome: | The proposed pipeline generates structured reports that enhance efficient, scalable legal analysis. |
Copied to clipboard
| Challenge: | Recent advances in deep learning have improved language generation systems, opening the door to improved forms of abstractive summarization. |
| Approach: | They propose to use neural encoder-decoder architectures to generate abstractive meeting summarizations that are particularly well-suited for multi-party conversation. |
| Outcome: | The proposed system could be used in a wide variety of real-world contexts, from business meetings to medical consultations to customer service calls. |
Copied to clipboard
| Challenge: | Existing methods that focus on learning a ranking across the whole candidate space are lacking user or task-specific training data. |
| Approach: | They propose an interactive ranking approach that actively selects pairs of candidates, from which the user selects the best. |
| Outcome: | The proposed approach outperforms existing methods in community question answering and extractive multidocument summarization and is an effective reward function for reinforcement learning. |
Copied to clipboard
| Challenge: | Existing studies have focused on disfluency detection and removal, with limited studies into its impact on downstream tasks. |
| Approach: | They propose to incorporate disfluency in summarization models to reduce the impact of replacement disfluencies on natural language processing tasks. |
| Outcome: | The proposed model improves on both public and real-life datasets and shows that it can handle disfluent data with up to 6.99-point degradation in Rouge-L score and replacement disfluencies have the highest negative impact. |
Copied to clipboard
| Challenge: | Large language models (LLMs) excel in various tasks, but often produce hallucinations . retrieved contexts, misrepresent information, or generate outright contradictions . |
| Approach: | They propose a framework that measures hallucination faithfulness of large language models . they introduce a leaderboard that leverages diverse human-annotated hallucinian examples . |
| Outcome: | The proposed framework improves hallucination evaluations by leveraging human-annotated examples. |
Copied to clipboard
| Challenge: | Existing approaches to interactive summarization are incomparable and divergent . a key gap in the development and adoption of interactive summaries is the lack of evaluation methodologies and benchmarks for meaningful comparison of systems. |
| Approach: | They propose an end-to-end evaluation framework for interactive summarization based on expansion-based interaction . framework includes procedure of collecting real user sessions, evaluation measures relying on summarizing standards, but adapted to reflect interaction. |
| Outcome: | The proposed evaluation framework is based on evaluations of baseline implementations and is available publicly as a benchmark. |
Copied to clipboard
| Challenge: | Recent advances in summarization are driven by the availability of large datasets such as the CNN-DailyMail corpus and the New York Times corpus. |
| Approach: | They propose a method for fine-tuning pretrained models for summarization in unsupervised manner . they use Wikipedia data to produce pseudo-summaries which contain characteristics of target dataset . |
| Outcome: | The proposed method achieves state-of-the-art, zero-shot abstractive summarization performance on CNN-DailyMail dataset and compares with other methods on other datasets. |
Copied to clipboard
| Challenge: | Hallucination is a well-known phenomenon in text generated by large language models . state-of-the-art LLMs still have a number of weaknesses, including the tendency to generate hallucinatory statements without considering the factuality . |
| Approach: | They propose a dataset that captures hallucinations made by retrieval-augmented LLMs . they propose to use these methods to help detect hallucinosity in QA tasks . |
| Outcome: | The proposed method captures hallucinations made by retrieval-augmented LLMs for QA tasks. |
Copied to clipboard
| Challenge: | Existing approaches to extract actionable suggestions from customer reviews are often mixed-intent, unstructured text. |
| Approach: | They propose a hybrid pipeline that uses a RoBERTa classifier and a precision–recall surrogate to extract actionable suggestions from customer reviews. |
| Outcome: | The proposed pipeline outperforms prompt-only, rule-based, and classifier-only baselines in extraction accuracy and cluster coherence. |
Copied to clipboard
| Challenge: | Empirically, we achieve the new state-of-the-art on all metrics (including human evaluation) on the CNN/Daily Mail dataset, as well as significantly higher abstractiveness scores. |
| Approach: | They propose a sentence-level policy gradient method that bridges computation between two neural networks in a hierarchical way while maintaining language fluency. |
| Outcome: | The proposed model achieves state-of-the-art on all metrics and higher abstractiveness scores on the CNN/Daily Mail dataset and faster training convergence than previous models. |
Copied to clipboard
| Challenge: | Applying natural language processing (NLP) techniques to the medical field is a prevailing trend nowadays and has great potential in many applications, such as key information extraction in medical literature. |
| Approach: | They propose to use a hierarchical encoder-tagger model to generate medical conversation summarization by identifying important utterances. |
| Outcome: | The proposed model outperforms baseline models and models and adds conversation-related features to improve performance. |
Copied to clipboard
| Challenge: | Recent advances on abstractive summarization have allowed substantial improvements in the quality of the model, but there is still scope for improvement. |
| Approach: | They propose novel multi-task architectures with high-level layer-specific sharing across multiple encoder and decoder layers of the three tasks and soft-sharing mechanisms. |
| Outcome: | The proposed model improves on the CNN/DailyMail and Gigaword datasets and on the DUC-2002 transfer setup. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have high computational costs and energy consumption, making their deployment in industrial settings difficult. |
| Approach: | They propose a small language model that compresses the embedding layer and reduces model size without significant loss of performance. |
| Outcome: | The proposed model reduces the embedding layer while maintaining performance while improving accuracy and performance. |
Copied to clipboard
| Challenge: | a dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications . identifying large, high-quality resources for summarization has called for creative solutions in the past. |
| Approach: | They present a summarization dataset of 1.3 million articles and summaries written by newsrooms of 38 major news publications. |
| Outcome: | The summarization dataset shows high diversity of summarizing styles . authors train existing methods on the data to evaluate its utility and challenges. |
Copied to clipboard
| Challenge: | Existing summarization models are limited in measuring the factual inconsistency of generated summaries. |
| Approach: | They propose a decoder overconfidence-regularizing objective as a hallucination risk measurement to better estimate the quality of generated summaries. |
| Outcome: | The proposed metric is reference-free and requires no training or modules . it records state-of-the-art correlation to human judgment on three sets of summary-quality annotations. |
Copied to clipboard
| Challenge: | Existing factuality-oriented abstractive summarization models only consider the integration of factual information and ignore the causes of factuual errors. |
| Approach: | They propose a factuality-oriented abstractive summarization model that can identify the causes of factual errors. |
| Outcome: | The proposed model outperforms state-of-the-art models in factual metrics. |
Copied to clipboard
| Challenge: | Existing work on Twitter uses extractive summarization to filter through information, but this approach often includes incomplete or redundant information. |
| Approach: | They propose to use Twitter data to generate 3100 gold-standard opinion summaries. |
| Outcome: | The proposed method outperforms previous work on extractive summarization models and fine-tunes to improve performance. |
Copied to clipboard
| Challenge: | Existing controllable summarization models do not allow users to specify their preference for a particular attribute of the generated summaries. |
| Approach: | They propose a novel training framework based on Constrained Markov Decision Process (CMDP) that includes a reward function and constraints to facilitate better summarization control. |
| Outcome: | The proposed model can be applied to control important attributes of summarization, including length, covered entities, and abstractiveness, while complying with a given attribute’s requirement. |
Copied to clipboard
| Challenge: | Adaptive training approaches do not consider the variation of learning difficulty in different training steps, making the learning deterministic and sub-optimal. |
| Approach: | They propose a dynamic token-level self-evolution training method that reweighs the training losses of different target tokens based on priors. |
| Outcome: | Empirically, the proposed method yields significant improvements on three translation tasks. |
Copied to clipboard
| Challenge: | a new generation of English-oriented Large Language Models significantly outperforms older LLMs on low-resource languages. |
| Approach: | They compare Bengali-oriented LLMs with open-weight and closed-source LLM models . they conclude that there is a need for a Bengali model, but lacks high-quality pretraining data . |
| Outcome: | The proposed model outperforms existing models on Bengali on low-resource languages . the results highlight biases in machine-translated datasets used for Bengali NLP tasks . |
Copied to clipboard
| Challenge: | Multilingual human preference data are difficult to obtain at scale, making it challenging to extend this framework to diverse languages. |
| Approach: | They propose a method where a reward model is trained on preference data in one source language and applied to other target languages. |
| Outcome: | The proposed approach is effective under comprehensive evaluation settings, including human evaluation. |
Copied to clipboard
| Challenge: | Recent studies have highlighted various neural metrics that align well with human evaluations. |
| Approach: | They propose a black-box adversarial framework that generates strong disagreements between human and victim evaluators. |
| Outcome: | The proposed framework can significantly improve the performance of human and victim evaluators. |
Copied to clipboard
| Challenge: | Existing methods for opinion summarization rely on human annotations, which may not be feasible. |
| Approach: | They propose to perform opinion summarization in an unsupervised manner by using a dictionary learning algorithm that implicitly captures semantic information from the review text. |
| Outcome: | The proposed algorithm performs well on SPACE and AMAZON datasets and performs controllable summarization to generate aspect-specific summaries using only a few samples. |
Copied to clipboard
| Challenge: | Existing LLMs require a new call to the inference endpoint/API for each new query . repeated calls to the endpoints/AP Is expensive and impractical for many real-world use cases. |
| Approach: | They compare the performance of various LLMs for query-based meeting summarization . they find that combining queries for the same context in a single prompt can be used to minimize repeated calls. |
| Outcome: | The proposed approach reduces the number of calls to the inference endpoints/APIs in meeting summarization tasks. |
Copied to clipboard
| Challenge: | Sequence-to-sequence (seq2sequ) models are a ubiquitous tool for text generation but they are not suitable for many other tasks. |
| Approach: | They propose to use UE techniques to identify out-of-domain (OOD) inputs where the model is susceptible to errors. |
| Outcome: | The proposed methods outperform heavyweight ensembles on the task of OOD detection. |
Copied to clipboard
| Challenge: | Sequence-to-sequence models have been used for natural language generation tasks such as machine translation and summarization. |
| Approach: | They propose to build a strong baseline based on general purpose sequence-to-sequence models for constituency parsing. |
| Outcome: | The proposed model outperforms existing models in natural language generation tasks without any explicit task-specific knowledge or architecture of constituent parsing. |
Copied to clipboard
| Challenge: | Structured data summarization involves generation of summaries from structured input data. |
| Approach: | They propose a hierarchical attention-based encoder-decoder model which leverages the structure in addition to the content of the tables. |
| Outcome: | The proposed model improves on the weathergov dataset by 30% over the current state-of-the-art. |
Copied to clipboard
| Challenge: | Abstractive summarization systems still suffer from faithfulness errors, authors say . prior work has proposed models that improve faithfulness, but it is unclear whether this improvement comes from an increased level of extractiveness of the outputs. |
| Approach: | They propose a faithfulness-abstractiveness trade-off curve that serves as a control . they also learn a selector to identify the most faithful and abstractive summary for a given document . |
| Outcome: | The proposed model achieves higher faithfulness scores while being abstractive than the baseline system on two datasets. |
Copied to clipboard
| Challenge: | Summarization studies work on increasing the scores that are given by automatic evaluation measures. |
| Approach: | They propose a simple but highly effective automatic evaluation measure of summarization, pruned Basic Elements. |
| Outcome: | The proposed measure outperforms ROUGE and BE in most cases and achieves highest correlation coefficient in TAC 2011 AESOP task. |
Copied to clipboard
| Challenge: | a novel approach to narrative event representation uses attention to re-contextualize events across the whole story . a recent study shows that attention is used to attach event semantics to tokens . |
| Approach: | They propose an unsupervised approach to narrative event representation using attention to re-contextualize events across the whole story. |
| Outcome: | The proposed approach achieves state of the art performance on multiple choice and story cloze tasks. |
Copied to clipboard
| Challenge: | Existing datasets for summarization of medical conversations are limited to conversation-summary pairs . a novel annotation framework is proposed to capture the summarizing process via an annotation task . |
| Approach: | They propose an incremental note generation framework that captures the human summarization process via an annotation task by instructing annotators to first incrementally create a draft note and polish it into a reference note. |
| Outcome: | The proposed framework shows that the human summarization process is much more efficient and accurate than the current method. |
Copied to clipboard
| Challenge: | Among recent NLP research, multi-document processing is gaining increasing attention due to the need to handle and process an increasing amount of textual data and available documents online. |
| Approach: | They propose to pre-train a generic multi-document model from a cross-document question answering pre-training objective by generating salient sentences from one document and challenging it to recover the sentence from which it was generated. |
| Outcome: | The proposed model outperforms zero-shot GPT-3.5 and GPT-4 in multiple document tasks and generates the correct answer and the salient sentence from a salient document. |
Copied to clipboard
| Challenge: | Existing research efforts to automate the document-to-slide generation process face a critical challenge: no publicly available dataset for training and benchmarking. |
| Approach: | They propose a dataset SciDuet that gathers papers and their corresponding slides from recent years’ NLP and ML conferences. |
| Outcome: | The proposed system outperforms state-of-the-art summarization baselines on both automated ROUGE metrics and qualitative human evaluation. |
Copied to clipboard
| Challenge: | Recent advances in natural language processing have impacted how models are trained for programming language tasks. |
| Approach: | They propose to use augmentation methods that yield consistent improvements in code translation and summarization by up to 6.9% and 7.5% respectively. |
| Outcome: | The proposed methods improve translation and summarization by 6.9% and 7.5% respectively. |
Copied to clipboard
| Challenge: | Prompt engineering is an essential technique for enhancing the abilities of large language models (LLMs) by providing explicit and specific instructions. |
| Approach: | They propose a new approach that uses text embeddings to obtain basis vectors by matrix decomposition and constructs a space for representing all prompts. |
| Outcome: | The proposed approach significantly outperforms state-of-the-art prompt paradigms on ten public reasoning benchmarks. |
Copied to clipboard
| Challenge: | Abstractive summarization is less prone to unfaithfulness issues than abstractive summaries . but, unfaitfulness problems, i.e., hallucinating new information, are still a problem in extractive summarisation . |
| Approach: | They propose a typology with five types of broad unfaithfulness problems that can appear in extractive summaries, including and beyond not-entailment. |
| Outcome: | The proposed metric shows that it detects unfaithful summaries faster than existing faithfulness evaluation metrics. |
Copied to clipboard
| Challenge: | Human evaluation is labor-intensive, expensive to scale, and difficult to design. |
| Approach: | They propose a set of guidelines for human evaluation of faithfulness in long-form summaries that address the following challenges: (1) How can we achieve high inter-annotator agreement on faithfulness scores? (2) How can our annotator minimize workload while maintaining accurate faithfulness? |
| Outcome: | The proposed framework reduces inter-annotator variance in faithfulness scores while minimizing annotator workload while maintaining accuracy. |
Copied to clipboard
| Challenge: | Existing studies on extractive summarization use finer-grained elementary discourse units . few studies exploited finer grained EDUs with little analysis and justification for the extractive unit selection . |
| Approach: | They propose an extractive model with Varying summary lengths that extracts fixed top-k salient sentences from the document as a summary. |
| Outcome: | The proposed model performs better on ROUGE scores than state-of-the-art models. |
Copied to clipboard
| Challenge: | Thematic progression is relevant to natural language processing applications dealing with discourse structure, argumentation structure, natural language generation, summarization and topic detection. |
| Approach: | They propose a toolkit for automatic analysis of thematic progression using a web interface. |
| Outcome: | ThemePro provides a visualization of the results including syntactic trees, hierarchical thematicity over propositions and thematic progression over whole texts. |
Copied to clipboard
| Challenge: | Existing studies focus on summarization and question-answering tasks, but neglect logical coherence within stories. |
| Approach: | They propose a model that leverages large language models to identify narrative gaps and generate coherent sentences that integrate seamlessly with the story’s emotional and logical flow. |
| Outcome: | The proposed model enhances narrative understanding and story generation, highlighting LLMs’ potential as effective logic checkers in story writing with logical coherence and emotional consistency. |
Copied to clipboard
| Challenge: | Query-focused tabular summarization is an emerging task in table-to-text generation . traditional transformer-based approaches face challenges due to token limitations and the complexity of reasoning over large tables. |
| Approach: | They propose a system that leverages tabular decomposition alongside a fine-tuned encoder-decoder model to improve summarization accuracy. |
| Outcome: | a new system outperforms the state-of-the-art REFACTOR model in a Query-focused tabular summarization task . the proposed system achieves a ROUGE-L score of 0.4437, outperforming the previous state- of-the art model . |
Copied to clipboard
| Challenge: | Transformer-based Language Models have become ubiquitous in natural language processing due to impressive performance on various tasks. |
| Approach: | They explore how sparsity affects network topology by exploiting mechanisms seen in biological networks . they show that model-agnostic sparsities are performant across diverse NLP tasks . |
| Outcome: | The proposed model-agnostic sparsity approaches are performant and efficient across NLP tasks. |
Copied to clipboard
| Challenge: | Existing methods for meeting summarization are limited and lack the robustness and context-based accuracy needed to maintain relevance. |
| Approach: | They propose a multi-LLM correction approach for meeting summarization using a two-phase process that mimics the human review process: mistake identification and summary refinement. |
| Outcome: | The proposed approach improves the quality of a given meeting summarization measured by relevance, informativeness, conciseness, and coherence. |
Copied to clipboard
| Challenge: | Existing models for diffusion generation are expensive and discrete, resulting in a large number of diffusion steps to generate text. |
| Approach: | They propose a text diffusion model that is fully non-autoregressive and employs a new form of self-conditioning and applies the diffusion process on the logit simplex space rather than the learned embedding space. |
| Outcome: | The proposed model outperforms state-of-the-art non-autoregressive models, requires fewer diffusion steps with minimal drop in performance, and is competitive with pretrained autoregressive sequence-to-sequence models. |
Copied to clipboard
| Challenge: | Existing document summarization methods focus on the text and filter out the non-textual content. Existing methods cannot meet the requirements of summarizing long text and multiple tables in each report. |
| Approach: | They propose a dataset for automatic document summarization that uses text and tabular data to produce a concise summary covering the input document's salient information. |
| Outcome: | The proposed method can produce a concise summary covering the input document's salient information. |
Copied to clipboard
| Challenge: | Existing studies on outline-conditioned text generation focus on generating text using provided outlines as rough sketches, but lack of clarity and rationality of the rough outlines hampers quality of the generated text. |
| Approach: | They propose a novel task that requires generating stories based on specific, sentence-level outlines. |
| Outcome: | The proposed framework improves the quality of precise outline-conditioned text generation. |
Copied to clipboard
| Challenge: | a new multilingual summarization model is being developed to help humanitarian experts process large amounts of secondary data to derive situational awareness and guide decision-making. |
| Approach: | They propose to use multilingual documents and annotated snippets to improve extraction of secondary data for humanitarian response experts. |
| Outcome: | The proposed model provides multilingual documents with informative snippets that have been annotated by humanitarian analysts over the past four years. |
Copied to clipboard
| Challenge: | Experimental results show that system summaries struggle to preserve syntactic meaning of source texts. |
| Approach: | They propose to incorporate syntactic information from source sentences into abstractive summaries by structure-infused copy mechanisms. |
| Outcome: | The proposed approach compares favorably to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing studies on cross-lingual summarization focus on pipeline methods or jointly training an end-to-end model through an auxiliary MT or MS objective. |
| Approach: | They propose a hierarchical model for the cross-lingual summarization task . the model is based on the conditional variational auto-encoder . |
| Outcome: | The proposed model generates better cross-lingual summaries than comparison models in the few-shot setting. |
Copied to clipboard
| Challenge: | Existing approaches for document dating assume accurate knowledge of document date, but this is not always available for arbitrary documents from the Web. |
| Approach: | They propose a Graph Convolutional Network (GCN) based document dating approach which exploits syntactic and temporal graph structures of document in a principled way. |
| Outcome: | The proposed approach outperforms state-of-the-art models on real-world datasets by 19% absolute accuracy points. |
Copied to clipboard
| Challenge: | Large language models excel in abstractive summarization tasks, delivering fluent and pertinent summaries. |
| Approach: | They conduct the first comprehensive study on context utilization and position bias in summarization. |
| Outcome: | The proposed benchmark compares two methods to alleviate position bias in summarization tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have advanced tasks like text summarization, but their size and computational demands limit their use in resource-constrained and privacy-centric settings. |
| Approach: | They propose a framework for distilling LLMs’ text summarization abilities into a compact, local model using a curriculum learning strategy that evolves from simple to complex tasks. |
| Outcome: | The proposed framework outperforms baseline models on CNN/DailyMail, XSum, and ClinicalTrial, and improves interpretability by providing insights into the summarization rationale. |
Copied to clipboard
| Challenge: | Submodular maximization with the greedy algorithm is an effective approach to extractive summarization. |
| Approach: | They propose a submodular maximization method that is 100 to 400 times faster than existing methods for extractive summarization. |
| Outcome: | The proposed method is 100 to 400 times faster than existing method based on integer-linear-programming formulations and achieves 95%-approximation. |
Copied to clipboard
| Challenge: | Existing methods for summarizing textual content are often ignored . relationshipal questions are ubiquitous and varied. |
| Approach: | They propose a method which generates a natural language summary of the relationship between two lexical items in a corpus without reference to a knowledge base. |
| Outcome: | The proposed method generates a natural language summary of the relationship between two lexical items in a corpus without reference to a knowledge base. |
Copied to clipboard
| Challenge: | Procedural text summarization task is a popular task in the NLP field because of its long length and complexity. |
| Approach: | They propose a procedural text summarization task with two granularity . they propose an Entity-State Graph-based Summarizer (ESGS) which aggregates contextual information for each procedure. |
| Outcome: | The proposed model can summarize the entire procedural text or give an overview for each step or both . Experiments on two datasets confirm the proposed model's effectiveness. |
Copied to clipboard
| Challenge: | Pre-trained language models have shown impressive results when fine-tuned on large summarization datasets. |
| Approach: | They analyze the training dynamics for generation models, focusing on summarization . they find that a propensity to copy the input is learned early in the training process . |
| Outcome: | The proposed model learns at different stages of fine-tuning, the authors show . they show that factual errors are learnt in later stages, but not at high-loss tokens . |
Copied to clipboard
| Challenge: | Argument Representation Coverage (ARC) assesses how well summaries preserve salient arguments . despite their fluency, LLMs frequently hallucinate or omit key content . |
| Approach: | They propose an evaluation framework that assesses how well summaries preserve salient arguments . they use argument representation coverage to distinguish between different information types . |
| Outcome: | The proposed framework assesses how well summaries preserve salient arguments . the authors show that LLMs capture some salient roles but omit critical information . |
Copied to clipboard
| Challenge: | a lack of high-quality English privacy policy corpus optimized for legal clarity and readability is limiting translation of privacy policies . 139 privacy policies are often considered "incomprehensible" due to technical jargon, legal language, and convoluted grammatical structures. |
| Approach: | They propose a high-quality English privacy policy corpus annotated by domain experts . they propose APPSI-139 to summarize and interpret privacy policies in English . |
| Outcome: | The proposed framework outperforms large language models in terms of readability and accuracy. |
Copied to clipboard
| Challenge: | Medical doctors spend 52 to 102 minutes per day writing clinical notes from patient encounters. |
| Approach: | They propose to use a new dataset to generate automated and manual clinical notes from doctor-patient conversations in a clinical setting. |
| Outcome: | The proposed model could reduce the time spent writing clinical notes from doctor-patient conversations in a clinical setting. |
Copied to clipboard
| Challenge: | Human evaluation captures quality but fails to capture diversity . statistical evaluation fails to catch models that plagiarize from training set . |
| Approach: | They propose a framework which evaluates both diversity and quality based on the optimal error rate of predicting whether a sentence is human-generated. |
| Outcome: | The proposed framework evaluates diversity and quality on summarization and chit-chat dialogue. |
Copied to clipboard
| Challenge: | Existing accuracy measures cannot evaluate the degree of personalization of summarization models. |
| Approach: | They propose to use a PENS dataset to analyze the degree of personalization of ten different summarization models. |
| Outcome: | The proposed measure can evaluate the degree of personalization of summarization models using the PENS dataset. |
Copied to clipboard
| Challenge: | Xu et al., 2006, show that model distillation can imbue efficient small language models with task-specific capabilities competitive with expensive teacher LLMs. |
| Approach: | They propose to distill outputs from a large teacher model to a small student model . they propose to use part-of-speech templates as higher-order linguistic features capable of capturing distinctive signals from teacher models that persist in distilled student outputs. |
| Outcome: | The proposed model distillation technique can imbue efficient small language models with task-specific capabilities competitive with (expensive) teacher LLMs. |
Copied to clipboard
| Challenge: | a quality summarization dataset requires the production and evaluation of summaries by trained humans and machines. |
| Approach: | They translate a summarization dataset in English and compare its performance to seven languages . they explore equivalence testing as an appropriate statistical paradigm for evaluating correlations between human and automated scoring of summaries . |
| Outcome: | The proposed method could be used in seven languages and compares performance across measures. |
Copied to clipboard
| Challenge: | a new study explores data manipulation techniques for improving abstractive summarization models without the need for any additional data. |
| Approach: | They propose a method of data synthesis with paraphrasing, data augmentation with sample mixing and curriculum learning with new difficulty metrics based on specificity and abstractiveness. |
| Outcome: | The proposed techniques improve abstractive summarization models without additional data . the proposed techniques can be applied in isolation and when combined . |
Copied to clipboard
| Challenge: | Existing text embedding benchmarks for financial domains are inadequately addressing the nuanced requirements of specialized domains like finance. |
| Approach: | They propose a finance-adapted embedding model that outperforms general-purpose models . they also introduce a new model, Fin-E5, which is also open-sourced . |
| Outcome: | The proposed framework outperforms general-purpose models on financial embedding tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are claimed to be capable of Natural Language Inference (NLI) |
| Approach: | They propose to use LLMs to probe their behavior using controlled experiments. |
| Outcome: | The proposed models perform significantly worse on NLI test samples which do not conform to these biases than those which do. |
Copied to clipboard
| Challenge: | Mental health disorders affect a significant portion of the global population . access to mental health support is limited in developing countries . |
| Approach: | They evaluated a 12-item descriptive MSE questionnaire and five well-known summarization models . they found that language models can generate coherent MSE summaries for doctors . |
| Outcome: | The proposed model can generate coherent summaries from MSEs in a conversational format. |
Copied to clipboard
| Challenge: | Existing methods to summarize text data are limited by the lack of data. |
| Approach: | They propose a method that uses external data to generate synthetic dialogues from short texts containing people and their interpersonal interactions. |
| Outcome: | The proposed method shows robust performance, generalizability, and scalability regardless of complexity of dialogues. |
Copied to clipboard
| Challenge: | a new text generation dataset is needed to controllable text summarization, but it lacks the domain knowledge. |
| Approach: | They propose to use existing text generation datasets to leverage input and control signals . they propose to annotate each meta-review sentence manually with a control signal . |
| Outcome: | The proposed method can be used to control the structure of a text generation dataset . it can be applied to a variety of tasks, including a task with a large number of meta-review sentences . |
Copied to clipboard
| Challenge: | Semantic parsing aims to transform natural language utterances into formal meaning representations (MRs) whereas an NL generator achieves the reverse, the two tasks are often studied separately. |
| Approach: | They propose a method of dual information maximization to regularize the learning process by matching the joint distributions of p and q of NLs. |
| Outcome: | The proposed method empirically maximizes the variational lower bounds of expected joint distributions of NL and MRs. |
Copied to clipboard
| Challenge: | Increasing use of virtual tutors has allowed for more efficient, personalized, and interactive AI-based learning experiences. |
| Approach: | They propose a task of Multi-modal Perspective based Dialogue Summarization (MM-PerSumm) that summarizes educational dialogues from three unique perspectives: the Student, the Tutor, and a Generic viewpoint. |
| Outcome: | The proposed model can summarize educational dialogues from three perspectives, while student-oriented summaries should distill learning points, track progress, and suggest scope for improvement. |
Copied to clipboard
| Challenge: | Existing methods to predict creation time of documents are based on time-stamp metadata, but none are available. |
| Approach: | They propose an attention-based neural document dating system which utilizes both context and temporal information in documents in a flexible and principled manner. |
| Outcome: | The proposed system outperforms neural and non-neural baselines on multiple real-world datasets. |
Copied to clipboard
| Challenge: | Abstractive summarization methods struggle with generating ungrammatical or even nonfactual contents. |
| Approach: | They evaluate ChatGPT's performance on extractive summarization and compare it with traditional fine-tuning methods on benchmark datasets. |
| Outcome: | The proposed pipeline performs better than abstractive methods on summary faithfulness and in-context learning. |
Copied to clipboard
| Challenge: | Recent studies have explored diffusion models for extractive summarization task, showcasing their remarkable capabilities. |
| Approach: | They propose a term-guided diffusion model for extractive summarization of legal documents that incorporates legal terminology into the model via a well-designed multifactor fusion noise weighting schedule. |
| Outcome: | The proposed model outperforms existing models on a self-constructed legal summarization dataset and achieves improvements of 3.10, 2.84, and 2.89 on three public datasets. |
Copied to clipboard
| Challenge: | a text fragment is discarded when it has a smaller context, causing it to acquire a new meaning or even become false. |
| Approach: | They build a dataset to study the effect of modifiers on the larger context . they focus on single-word modifiers, the smallest unit that can be considered disposable . |
| Outcome: | The proposed dataset aims to determine whether modifiers can be removed without undesirable consequences. |
Copied to clipboard
| Challenge: | Recent work has shown that small, context-dependent shifts in word distributions can be used to apply and detect watermarks, but little work has analyzed the impact of these perturbations on the quality of generated texts. |
| Approach: | They propose a framework that allows for analysis of the impact of watermark settings on the quality of generated texts. |
| Outcome: | The proposed framework provides easy visualization of the quality-detection trade-off of watermark settings. |
Copied to clipboard
| Challenge: | Abstractive summarization systems are difficult to perform due to the unavailability of the parallel data for low-resource languages like Bengali. |
| Approach: | They propose a graph-based unsupervised abstractive summarization system in Bengali text documents that requires only a Part-Of-Speech (POS) tagger and a pre-trained language model trained on Bengali texts. |
| Outcome: | The proposed system outperforms baselines without human-annotated reference summaries on a human-random dataset with Bengali text. |
Copied to clipboard
| Challenge: | Existing benchmarks for summarization quality evaluation lack diverse input scenarios, focus on narrowly defined dimensions, and struggle with subjective and coarse-grained annotation schemes. |
| Approach: | They propose to use AI to help human annotations and identifie potentially hallucinogenic input texts. |
| Outcome: | The proposed benchmarks improve on existing benchmarks in terms of input diversity, granularity of human annotations, and evaluation dimensions. |
Copied to clipboard
| Challenge: | Biomedical data and benchmarks are highly valuable but limited in low-resource languages such as English. |
| Approach: | They propose a translation model in Vietnamese that trains a pretrained Encoder-Decoder Transformer model on 20 million translated abstracts. |
| Outcome: | The proposed model can translate and produce both pretrained and supervised biomedical data in two biomedically important domains. |
Copied to clipboard
| Challenge: | Existing studies have shown that large language models contain linguistic and societal biases, but it is unclear how these biase amplify to downstream tasks. |
| Approach: | They investigate how name-nationality bias propagates from pre-training to downstream tasks . they show that these biases manifest themselves as hallucinations in summarization . |
| Outcome: | The proposed model can reduce the rate of hallucinations, but does not change the types of biases that do appear. |
Copied to clipboard
| Challenge: | Summarization quality evaluation is a non-trivial task in text summarization. |
| Approach: | They propose a unified multi-scenario summarization evaluation model that shares cross-sceenario knowledge and uses a self-supervised training paradigm to optimize the model without extra human labeling. |
| Outcome: | The proposed model can achieve comparable performance with existing methods for three evaluation scenarios. |
Copied to clipboard
| Challenge: | Existing methods for controllable summarization fail to generate entity-centric summaries. |
| Approach: | They propose to use a human-annotated data set EntSUM to generate controllable summarization with a focus on named entities as the aspects to control. |
| Outcome: | The proposed data set shows that existing methods fail to generate entity-centric summaries. |
Copied to clipboard
| Challenge: | Movie screenplay summarization requires an understanding of long input contexts and elements unique to movies. |
| Approach: | They propose a dataset for movie screenplay summarization that includes movie screenplayers accompanied by their Wikipedia plot summaries. |
| Outcome: | The proposed dataset includes 2200 movie screenplays accompanied by their Wikipedia plot summaries. |
Copied to clipboard
| Challenge: | Existing summarization systems for multilingual text summarizing are limited due to the lack of large-scale data in multiple languages. |
| Approach: | They propose a multilingual summarization system that can understand documents in multiple languages and generate summaries in the corresponding language. |
| Outcome: | The proposed model improves over monolingual models in all languages and transferable to other languages. |
Copied to clipboard
| Challenge: | Query-focused summarization has been considered as an important extension for text summarizing . lack of large-scale datasets hinders its development . |
| Approach: | They propose to integrate text summarization and question answering into a prefix-based pretraining strategy for few-shot learning in query-focused summarizing. |
| Outcome: | The proposed prefix-based pretraining outperforms fine-tuning on query-focused summarization. |
Copied to clipboard
| Challenge: | Recent advances on models and metrics should benefit and inform each other, authors argue . bidimensional leaderboards allow for fast, accurate evaluation of language generation models . |
| Approach: | They propose a bidimensional leaderboard that tracks progress in language generation models and metrics for their evaluation. |
| Outcome: | The proposed leaderboards track progress in language generation models and metrics for their evaluation. |
Copied to clipboard
| Challenge: | Abstractive summarization methods suffer from inferior performance compared to extractive methods. |
| Approach: | They propose a reddit TIFU dataset and a new abstractive summarization model . they use multi-level memory networks to store information from different levels of abstraction . |
| Outcome: | The proposed model outperforms state-of-the-art summarization models with multi-level memory . the proposed dataset is highly abstractive and outperformed existing models with the proposed model . |
Copied to clipboard
| Challenge: | Existing theories claim that pretraining models learn linguistic knowledge from the pretraining corpus, but scientific explanations for these benefits remain unknown. |
| Approach: | They propose to use random character n-grams to test models on real corpora to see if the small residual benefit of using real data could be accounted for by the structure of the pretraining task. |
| Outcome: | The proposed task performs on documents consisting of character n-grams, whereas pretrained models perform on real corpora with no residual benefit. |
Copied to clipboard
| Challenge: | Existing models for summarization of legal documents rely on external knowledge to generate abstracts. |
| Approach: | They propose an entity-driven approach that learns the model to generate factual hallucinations . they evaluate legal documents in English and French to evaluate their results . |
| Outcome: | The proposed approach reduces non-factual hallucinations and maximizes summary coverage and factual hallucines at entity-level. |
Copied to clipboard
| Challenge: | DiscoScore is a parametrized discourse metric that uses BERT to model discourse coherence . it is weak when operated at system level, and is therefore not reliable in a way to spot improvements . |
| Approach: | They propose a parametrized discourse metric which uses BERT to model discourse coherence from different perspectives. |
| Outcome: | The proposed model outperforms existing models on document-level machine translation and summarization. |
Copied to clipboard
| Challenge: | ChatGPT and GPT-4 are popular as evaluation metric for complex generative tasks . however, they are not ready as human replacements due to significant limitations . |
| Approach: | They conduct extensive analysis to examine the stability and reliability of LLMs as automatic evaluators for abstractive summarization. |
| Outcome: | The proposed methods outperform the commonly used automatic metrics but are not ready for human evaluation due to significant limitations. |
Copied to clipboard
| Challenge: | Recent research in mechanistic interpretability has revealed that Large Language models contain disentangled, human-understandable components. |
| Approach: | They propose a framework that first identifies causal task features through frequency recall and interventional filtering, then selects “Feature-Resonant Data” that maximally activates task features for fine-tuning. |
| Outcome: | The proposed framework outperforms existing models on mathematical reasoning, summarization, and translation tasks while using only 50% of the data. |
Copied to clipboard
| Challenge: | Existing studies on domain generalization have sophisticated training algorithms. |
| Approach: | They propose a lightweight, weight averaging approach to domain generalization for abstractive summarization using prefix tuning and weight adjusting. |
| Outcome: | The proposed method performs better on four diverse summarization domains compared to baselines. |
Copied to clipboard
| Challenge: | Existing studies on content importance do not consider semantics and context when evaluating importance. |
| Approach: | They apply information theory to pre-trained language models to define the concept of importance from the perspective of information amount. |
| Outcome: | Experiments on CNN/Daily Mail and New York Times show that the proposed model can model the importance of content better than previous methods based on F1 and ROUGE scores. |
Copied to clipboard
| Challenge: | Abstractive document summarization models are often trained on limited supervised data . authors present three objectives for pretraining abstractive summarizing models . |
| Approach: | They propose to pre-train a SEQ2SEQ based abstractive summarization model on unlabeled text. |
| Outcome: | The proposed method improves on two benchmark summarization datasets with 19GB of text . the goal is sentence reordering, next sentence generation and masked document generation . |
Copied to clipboard
| Challenge: | Abstractive summarization models have seen great improvements in recent years, but there is limited understanding of the strategies different models employ and how they relate their understanding of language. |
| Approach: | They characterize how one popular abstractive model uses an explicit copy/generation switch to control its level of abstraction vs extraction . they find that abstractive summarization models lack the semantic understanding necessary to generate paraphrases that are both abstractive and faithful to the source document. |
| Outcome: | The proposed model uses syntactic boundaries to truncate sentences that are often copied verbatim. |
Copied to clipboard
| Challenge: | Existing models generate fluent and coherent summaries, but inconsistencies can be found in generated summary. |
| Approach: | They propose to use symbolic knowledge distillation to improve the factual consistency of smaller pretrained models for dialogue summarization. |
| Outcome: | The proposed model outperforms baseline models in BART, PEGASUS, and Flan-T5 in factual consistency and accuracy. |
Copied to clipboard
| Challenge: | Existing work has addressed each element individually, but this study focuses on LipKey, the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles. |
| Approach: | They propose a novel news dataset that consists of highly absent keyphrases . they combine lips keyphrase and TF-IDF to obtain abstractive summaries . |
| Outcome: | The proposed dataset is the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles. |
Copied to clipboard
| Challenge: | IR for precision medicine often involves looking for multiple pieces of evidence that characterize a patient case. |
| Approach: | They propose a document reranking approach that combines neural query-document matching and text summarization toward such retrieval scenarios. |
| Outcome: | The proposed approach achieves state-of-the-art performance on NIST's TREC-PM track dataset. |
Copied to clipboard
| Challenge: | Contemporary leading-edge systems for abstractive (long) text summarization employ Transformer encoderdecoder architectures that only consider the nuclearity annotation . |
| Approach: | They propose to incorporate Rhetorical Structure Theory into a novel summarization model that incorporates both the types and uncertainty of rhetorical relations. |
| Outcome: | The proposed model outperforms state-of-the-art models on automatic metrics and human evaluation. |
Copied to clipboard
| Challenge: | Generating diverse sequences exhibit semantically one-to-many relationships between source and target sequences. |
| Approach: | They propose to separate diversification from generation using a general plug-and-play module that wraps around and guides an existing encoder-decoder model. |
| Outcome: | The proposed method shows that diversification and generation are separate steps in the same model and that the model is robust. |
Copied to clipboard
| Challenge: | Using a variety of language generation models, ensembling models is challenging during inference. |
| Approach: | They propose a method that decodes text models that do not assume a shared vocabulary, tokenization or generation order. |
| Outcome: | The proposed method outperforms models decoded in isolation over various scenarios. |
Copied to clipboard
| Challenge: | Recent advances in text autoencoders have significantly improved the quality of the latent space, allowing models to generate consistent text from aggregated latent vectors. |
| Approach: | They develop a framework which searches input-output word overlap for latent vector aggregation. |
| Outcome: | The proposed framework improves the quality of the latent space and establishes state-of-the-art performance on two opinion summarization benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for summarizing documents are inconsistent due to the difficulty of manual evaluation. |
| Approach: | They propose a method where summaries are evaluated by multiple annotators against the source document via manually highlighted salient content. |
| Outcome: | The proposed method improves inter-annotator agreement while highlighting differences among systems. |
Copied to clipboard
| Challenge: | Existing solutions focus on efficient attentions or divide-and-conquer strategies, but these methods sacrifice global context, leading to incoherent and uninformative summaries. |
| Approach: | They propose to leverage the memory-efficient nature of divide-and-conquer methods while preserving global context. |
| Outcome: | The proposed framework improves informativeness, faithfulness, and coherence over baselines on government reports, meeting transcripts, screenplays, scientific papers, and novels. |
Copied to clipboard
| Challenge: | Recent advances in abstractive summarization systems produce factually inconsistent text . this is emphasized in tasks like summarizing, which often produce inconsistent text with no input article . |
| Approach: | They use reinforcement learning to optimize for factual consistency and explore trade-offs . they use textual-entailment rewards to optimize the accuracy of the generated summaries . |
| Outcome: | The proposed method improves faithfulness, salience and conciseness of the generated summaries. |
Copied to clipboard
| Challenge: | Existing contrastive decoding methods that handle conflict lack adaptability and can degrade performance in low conflict settings. |
| Approach: | They propose a token-level algorithm for principled conflict resolution and enhanced faithfulness that resolves conflict by utilizing confidence-aware measures and the generalized divergence between parametric and contextual distributions. |
| Outcome: | The proposed algorithm achieves 9.2 points on average in QA, summarization, and long-form question answering (LFQA) benchmarks and improves factuality by 2.5 points on the key benchmarks. |
Copied to clipboard
| Challenge: | Existing models for summarizing medical conversations do not take clinical knowledge into account and are difficult to control. |
| Approach: | They propose a transformer-based sequence-to-sequence architecture for summarizing medical conversations by integrating medical domain knowledge from the Unified Medical Language System (UMLS). |
| Outcome: | The proposed model achieves state-of-the-art ROUGE score improvements of 0.8-2.1 points (including 6.2% error reduction in the PE section) it incorporates medical domain knowledge from the Unified Medical Language System (UMLS). |
Copied to clipboard
| Challenge: | Fine-tuning is the prevalent paradigm for using large pretrained language models for downstream tasks, but it requires updating and storing all the parameters of the LM. |
| Approach: | They propose a lightweight alternative to fine-tuning for natural language generation tasks that optimizes a sequence of continuous vectors, which they call the prefix. |
| Outcome: | The proposed approach outperforms fine-tuning in the full data setting and extrapolates better to examples with topics that are unseen during training. |
Copied to clipboard
| Challenge: | Pretrained large language models can reproduce harmful social biases in constrained settings, such as summarization. |
| Approach: | They propose a method to generate input documents with carefully controlled demographic attributes and then apply it to a controlled setting. |
| Outcome: | The proposed method allows to generate input documents with carefully controlled demographic attributes while working with real-world input documents. |
Copied to clipboard
| Challenge: | Neural text generation has been quite successful recently, but during training time, only one reference is considered for each example, even though there are often multiple references available. |
| Approach: | They propose an algorithm to generate exponentially many pseudo-references by compressing existing references into lattices and traversing them to generate new pseudo-References. |
| Outcome: | The proposed model significantly improves on baselines in machine translation and image captioning, and is comparable to existing models. |
Copied to clipboard
| Challenge: | Existing pretrained models require domain-specific additional information to be effective. |
| Approach: | They propose a pre-trained model for multi-document representation with a focus on summarization that uses efficient encoder-decoder transformers to simplify the processing of concatenated input documents. |
| Outcome: | PRIMERA outperforms current state-of-the-art models on most datasets with large margins . PRImerA uses efficient encoder-decoder transformers to simplify processing of concatenated input documents. |
Copied to clipboard
| Challenge: | Detecting contradictions in texts is often regarded as determining relation between hypothesis and piece of premise. |
| Approach: | They propose a human-annotated dataset to study self-contradictions in long documents . they analyze the capabilities of four open-source and commercially available LLMs . |
| Outcome: | The proposed dataset outperforms open-source LLMs on document-level tasks but struggles with self-contradictions that require more nuance and context. |
Copied to clipboard
| Challenge: | Existing statistical phrasal or hierarchical machine translation systems relies on a large set of translation rules which results in engineering challenges. |
| Approach: | They propose to use factorized grammar from the field of linguistics as more general translation rules from XTAG English Grammar to generate a manually crafted summarization dataset. |
| Outcome: | The proposed method outperforms existing methods on low-resource language translation tasks with less training data. |
Copied to clipboard
| Challenge: | Existing models for dialogue summarization focus on extracting the main events of short conversations, but real-world dialogues are difficult to train. |
| Approach: | They propose three strategies to deal with the lengthy input problem and locate relevant information using long dialogue datasets. |
| Outcome: | The retrieve-then-summarize pipeline models yield the best performance on three long dialogue datasets. |
Copied to clipboard
| Challenge: | Recent studies have raised concerns about the potential threats large language models pose to academic integrity and copyright protection. |
| Approach: | They propose a dataset of 46.5K synthetic text pairs that represent three major types of plagiarism: verbatim copying, paraphrasing, and summarization. |
| Outcome: | The proposed dataset shows that GPT-3.5 Turbo can produce high-quality paraphrases and summaries without significantly increasing text complexity compared to GPT-4 Turbo. |
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) can grasp the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent. |
| Approach: | They propose a framework for multimodal large language models to grasp the intention of a question and decompose it into a series of visual recognition sub-tasks to find out the answer. |
| Outcome: | The proposed framework improves the accuracy of complex video-related questions by 29.6% and 17.2% on CVQA and the existing VQA datasets. |
Copied to clipboard
| Challenge: | Statistical language modeling and translation with transformers have found many successful applications in program understanding and generation tasks. |
| Approach: | They propose an architecture-independent approach for leveraging syntactic hierarchies of source code . they use syntax trees to extract syntak hierarchical structures and integrate them into context window . |
| Outcome: | The proposed approach achieves state-of-the-art in code completion and summarization for Python in the CodeXGLUE benchmark. |
Copied to clipboard
| Challenge: | Multi-document summarization (MDS) aims at combining information spread across multiple documents . a single document often covers the full summary content . |
| Approach: | They propose a measure to evaluate the degree to which a summary is "disperse" they propose to combine information from multiple documents into a single document to generate a concise summary . |
| Outcome: | The proposed measure evaluates the degree to which a summary is "disperse" the measure is applied to several popular MDS datasets and state-of-the-art systems. |
Copied to clipboard
| Challenge: | Abstractive summarization quality has been improved but there is a lack of data for conversation summarizing applications. |
| Approach: | They propose to build a conversation summarization dataset with human written summaries from internet forums. |
| Outcome: | The proposed dataset can be easily expanded to improve conversation summarization applications. |
Copied to clipboard
| Challenge: | Existing evaluation metrics do not capture meeting-specific errors, leading to ineffective assessment. |
| Approach: | They examine the relationship between established metrics and human evaluations to determine what challenges and errors are captured by correlating metric scores with human evaluation. |
| Outcome: | The proposed measures show weak correlations with human evaluations and a third of the correlations show error masking. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) generate misleading or outright incorrect information. |
| Approach: | They propose a method that debiases uncertainty scores on output length and uses residuals as corrected, length-invariant estimates. |
| Outcome: | The proposed method improves over nominally length-normalized methods on machine translation, summarization, and question-answering tasks. |
Copied to clipboard
| Challenge: | Large language models exhibit positional bias, a problem that can undermine the completeness of conversation summarizations. |
| Approach: | They propose a semantic similarity-based sentence-level metric to quantify positional bias in conversational summaries. |
| Outcome: | The proposed benchmark provides the first systematic evaluation of positional bias in conversational summarization across languages and contexts. |
Copied to clipboard
| Challenge: | Existing methods for generating comparative summaries that highlight similarities and contradictions in input documents are lacking large parallel training data for their training. |
| Approach: | They propose a method for generating comparative summaries that highlight similarities and contradictions in input documents by using a neural interpretation of traditional concept-to-text generation systems. |
| Outcome: | The proposed model is compared with conventional methods in the domain of nutrition and health, where the existing models lack large parallel training data. |
Copied to clipboard
| Challenge: | Goal-oriented conversations often have sub-dialogue structure, but it can be domain-dependent . Increasingly, language understanding applications involve conversational speech and text . |
| Approach: | They propose an unsupervised approach to learning hierarchical conversation structure . they use turn and sub-dialogue segment labels to decode the structure based on dialogue acts and subtasks . |
| Outcome: | The proposed approach improves neural models for three conversation-level understanding tasks. |
Copied to clipboard
| Challenge: | Existing literature review models have addressed literature review generation, but lack of large-scale datasets has been a stumbling block. |
| Approach: | They propose to use a large-scale dataset to evaluate automatic literature review generation models. |
| Outcome: | The proposed model can generate summaries comparable to human-written reviews while lacking detailed information. |
Copied to clipboard
| Challenge: | Existing methods for summarizing text have not captured the salient information from an article. |
| Approach: | They propose a table-guided abstractive biography summarization that utilizes factual tables to capture important information and generate a summary of a biography. |
| Outcome: | The proposed method is the first large-scale biography summarization dataset with tables. |
Copied to clipboard
| Challenge: | Existing supervised fine-tuning (SFT) methods focus on directly generating the target output without leveraging the benefits of intermediate steps or initial guidance. |
| Approach: | They propose a task-agnostic framework that enables models to generate intermediate "warmup" sequences that are iteratively refined to maximize their contribution to the final output. |
| Outcome: | The proposed framework outperforms traditional supervised fine-tuning methods on translation, summarization, and multi-choice question answering tasks. |
Copied to clipboard
| Challenge: | a large corpus of documents is available for summarization tasks in English . supervised methods require adequate corpora for summarizing . |
| Approach: | They describe a corpus of catalan and spanish newspapers that can be used to train summarization models for Catalan, Spanish and other languages. |
| Outcome: | The proposed corpus can be used to train summarization models for Catalan and Spanish. |
Copied to clipboard
| Challenge: | Existing studies focus on sentence-level inference, which limits its application in downstream NLP problems. |
| Approach: | They propose to construct a large-scale dataset for document-level NLI that can be used to study NLP problems. |
| Outcome: | The proposed model performs well on popular sentence-level benchmarks and generalizes well to out-of-domain NLP tasks that rely on inference at document granularity. |
Copied to clipboard
| Challenge: | Existing models do not provide an efficient way to locate information that enters the common ground. |
| Approach: | They propose a method based on segmentation of a conversation into themes followed by their summarization and obtain the location of information transfers by computing the distance between the theme summary and the different utterances produced by a speaker. |
| Outcome: | The proposed method is based on the segmentation of a conversation into themes followed by their summarization and obtains the location of information transfers by computing the distance between the theme summary and the different utterances produced by a speaker. |
Copied to clipboard
| Challenge: | Existing summarization methods read through document only once to generate a document representation, resulting in a sub-optimal representation. |
| Approach: | They propose an iterative model for supervised extractive text summarization which polishes the document representation on many passes through the document. |
| Outcome: | The proposed model outperforms state-of-the-art extractive systems on CNN/DailyMail and DUC2002 datasets. |
Copied to clipboard
| Challenge: | lexical overlap is a common evaluation metric for extractive summarization, but recent studies reveal its limitations. |
| Approach: | They propose a facet-aware evaluation setup for better assessment of information coverage in extractive summaries. |
| Outcome: | The proposed evaluation setup improves human correlation with extractive summarization datasets and improves comparative analysis. |
Copied to clipboard
| Challenge: | Existing methods for generating generic summarizations can't be used to generalize to these domains without seeing in-domain training data. |
| Approach: | They use a dataset of real-world aspect-oriented summaries to annotate articles from two different news sub-domains. |
| Outcome: | The proposed approach produces better focused summaries than existing systems without seeing in-domain training data. |
Copied to clipboard
| Challenge: | Existing models for dialog generation are challenging to train using the standard Seq2Seq models. |
| Approach: | They propose a framework for Hierarchical Transformer Encoders that can be morphed into any hierarchical transformer by using specially designed attention masks and positional encodings. |
| Outcome: | The proposed framework can be morphed into any hierarchical encoder, including HRED and HIBERT like models, by using specially designed attention masks and positional encodings. |
Copied to clipboard
| Challenge: | Recent studies have shown that sentence-based extractive models result in redundant or uninformative phrases in the extracted summaries. |
| Approach: | They propose a discourse-aware neural summarization model that extracts sub-sentential discourse units as candidates for extractive selection on a finer granularity. |
| Outcome: | Experiments show that the proposed model outperforms state-of-the-art models on popular summarization benchmarks. |
Copied to clipboard
| Challenge: | Sentence summarization systems that use latent space to reconstruct the source sentence are unwillingly exploited. |
| Approach: | They propose a method that uses language modeling and semantic similarity metrics to find a high-scoring summary. |
| Outcome: | The proposed method achieves state-of-the-art for unsupervised sentence summarization according to ROUGE scores. |
Copied to clipboard
| Challenge: | a new task is needed to distinguish between foreground and background events in news articles . |
| Approach: | They propose a task of distinguishing between foreground and background events in news articles . they also identify the general temporal position of background events relative to the foregoing period . |
| Outcome: | The proposed model achieves good performance on a dataset of news articles . |
Copied to clipboard
| Challenge: | Existing metrics for conditional natural language generation rely on pairwise comparisons between a single generated text and the best-matching reference. |
| Approach: | They propose a family of meta-metrics that build on existing pairwise distance functions to evaluate conditional natural language generation models. |
| Outcome: | The proposed method evaluates the ability of a model to generate text matching diversity in references in visual description and summarization. |
Copied to clipboard
| Challenge: | Existing methods for query-focused table summarization struggle with complex reasoning and token-limit issues. |
| Approach: | They propose a Fast, Accurate, and Privacy-Compliant table summarization approach via Offline Template Generation. |
| Outcome: | The proposed method outperforms baseline methods on widely-used benchmarks. |
Copied to clipboard
| Challenge: | Scientific literature review generation aims to extract and organize important information from an abundant collection of reference papers and produces corresponding reviews while lacking a clear and logical hierarchy. |
| Approach: | They propose a task to generate a hierarchical catalogue of a review paper given various references by using a database of 7.6k literature review catalogues and 389k reference papers. |
| Outcome: | The proposed method produces a hierarchical catalogue of a review paper given various references. |
Copied to clipboard
| Challenge: | Existing methods for summarizing arguments are incapable of distinguishing between generated key points of different qualities. |
| Approach: | They propose an extractive approach that generates concise, high quality key points . they propose to use a clustering approach to generate key points from raw arguments . |
| Outcome: | The proposed method outperforms state-of-the-art methods for key point generation . it offers concise, high quality generated key points with higher coverage of reference summaries . |
Copied to clipboard
| Challenge: | Current summarization systems only produce plain, factual headlines, far from the practical needs for exposure and memorableness of the articles. |
| Approach: | They propose a task to generate relevant headlines with three style options . they propose combining summarization and reconstruction tasks into a multitasking framework . |
| Outcome: | The proposed method outperforms the state-of-the-art summarization model by 9.68% . it can generate relevant, fluent headlines with humor, romance and clickbait . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) often experience “contextual hallucination” where they prioritize self-generated content over input context, leading to a disregard for pertinent details. |
| Approach: | They propose a method that dynamically adjusts attention maps to enhance contextual relevance by using a trained classifier to identify attention maps likely to induce hallucinations. |
| Outcome: | The proposed approach reduces hallucinations across open-source models on summarization and open-book QA tasks. |
Copied to clipboard
| Challenge: | Existing methods for dialogue summarization only apply to specific scenarios and domains. |
| Approach: | They propose a pre-trained model specifically designed for multi-scenario multi-domain dialogue summarization. |
| Outcome: | The proposed model significantly outperforms state-of-the-art models on three dialogue summarization datasets from different scenarios and domains. |
Copied to clipboard
| Challenge: | Recent developments in balancing usefulness and safety of large language models raise a critical question . current attacks, especially adversarial ones that manipulate malicious prompts, often aim to manipulate the input . |
| Approach: | They show that LLMs can effectively summarize malicious long documents but often refuse to translate them. |
| Outcome: | The findings highlight a vulnerability in LLMs that can't translate or summarize documents . the study focuses on LLM models, Gemini and GPT-4, which can' be exploited . |
Copied to clipboard
| Challenge: | Recent work on opinion summarization has focused on extracting fragments from reviews, but we use novel sentences to generate abstractive summaries. |
| Approach: | They propose an abstractive summarizer which does not use summaries in training and is trained end-to-end on a large collection of reviews. |
| Outcome: | The proposed model produces fluent and coherent summaries reflecting consensus opinions on Amazon and Yelp reviews. |
Copied to clipboard
| Challenge: | Existing Persian datasets are small and lack content diversity . lack of high-quality data has slowed development of NLP models and open-source LLMs for Persian. |
| Approach: | They propose a Persian dataset of 72.9B tokens that is preprocessed and deduplicated to ensure high data quality. |
| Outcome: | The proposed model performs well on key Persian NLP tasks. |
Copied to clipboard
| Challenge: | Existing studies have focused on examining hallucinations stemming from static input, such as in summarization or machine translation. |
| Approach: | They propose a knowledge-augmented generator that produces information that remains grounded in contextual knowledge regardless of alterations in the context. |
| Outcome: | The proposed method is designed to produce information that remains grounded in contextual knowledge, regardless of alterations in the context. |
Copied to clipboard
| Challenge: | Existing research on news summarization focuses on single-language single-document (SLSD), single-linguistic multi-document or cross-language multi-doc (CLSD) however, in real-world scenarios, news articles often involve multiple documents in different languages, i.e., mixed-language MLMD. |
| Approach: | They propose a mixed-language multi-document news summarization dataset with four different languages and 10,992 source document cluster and target summary pairs. |
| Outcome: | The proposed dataset contains four different languages and 10,992 source document cluster and target summary pairs. |
Copied to clipboard
| Challenge: | Existing LLMs are limited by text-context budgets, resulting in token-expensive storage of raw trajectories . Optical Context Retrieval Memory (OCR-Memory) renders historical tra-jectorios into images annotated with unique visual identifiers. |
| Approach: | They propose a framework that leverages the visual modality as a high-density representation of agent experience. |
| Outcome: | Optical Context Retrieval Memory (OCRM) renders historical trajectories into images annotated with unique visual identifiers. |
Copied to clipboard
| Challenge: | Recent studies have shown that current models are prone to generating unfaithful summaries . a proposed method is effective in identifying and correcting extrinsic hallucinations . |
| Approach: | They propose a model-agnostic post-processing technique to correct unfaithful summaries . they generate alternative candidates where names and quantities are replaced with compatible ones . |
| Outcome: | The proposed method corrects extrinsic hallucinations in unfaithful summaries. |
Copied to clipboard
| Challenge: | Existing multimodal summarization models ignore the contribution of visual modalities . we propose a novel contribution network to consider different contributions of images . |
| Approach: | They propose a Coarse-to-Fine contribution network for multimodal summarization to consider different contributions of images for summarizing. |
| Outcome: | The proposed system outperforms baselines on the visual and textual modalities. |
Copied to clipboard
| Challenge: | Abstractive summarization models produce factually inconsistent summaries that are not supported by the original article. |
| Approach: | They propose a fact-aware filtering mechanism that improves the factuality of abstractive summarization models. |
| Outcome: | The proposed method improves the quality of training data and the factuality of generated summaries. |
Copied to clipboard
| Challenge: | Recent studies have found that summaries generated by large language models (LLMs) are favored by human annotators when compared to reference summary from widely used summarization datasets. |
| Approach: | They propose to use large language models (LLMs) as reference learning settings for smaller text summarization models to investigate whether their performance can be substantially improved. |
| Outcome: | The proposed model outperforms standard supervised fine-tuning and human evaluations while retaining human-level performance. |
Copied to clipboard
| Challenge: | Context information is one of the key factors for extractive summarization, but other factors can be used to identify sentence importance. |
| Approach: | They propose to disentangle context and pattern factors for extractive summarization . they separate context and patterns for a better generalization ability in low-resource setting . |
| Outcome: | The proposed model can be used in the zero-shot setting or fine-tuned in the few-shot settings. |
Copied to clipboard
| Challenge: | Existing text summarization models lack guiding entities to ensure that entities are present in summaries. |
| Approach: | They propose a controllable abstractive sentence summarization model which generates summaries with guiding entities. |
| Outcome: | The proposed model outperforms the state-of-the-art models in evaluation scores and informativeness metrics. |
Copied to clipboard
| Challenge: | Efficient document summarization requires evaluation measures that can rank a set of systems based on an average score and highlight which individual summary is better than another. |
| Approach: | They propose a hybrid evaluation measure for document summarization called HOLMS that combines both language models pre-trained on large corpora and lexical similarity measures. |
| Outcome: | The proposed measure outperforms ROUGE and BLEU on several extractive summarization datasets for both linguistic quality and pyramid scores. |
Copied to clipboard
| Challenge: | Developing social summarization systems is becoming more and more critical . but, the publicly available and high-quality large scale social summaries are rare . |
| Approach: | They propose to build a social summarization dataset using twitter's hot events . they collect user relations, hashtags and user profiles to evaluate their summarizing methods . |
| Outcome: | The proposed dataset is based on a dataset from twitter with 12 real world hot events with 44,034 tweets and 11,240 users. |
Copied to clipboard
| Challenge: | Generating concise summaries of news events is a challenging task for newcomers to a news story. |
| Approach: | They propose a task of background news summarization that complements each timeline update with a background summary of relevant preceding events. |
| Outcome: | The proposed system performs well on a question-answering-based evaluation metric, Background Utility Score (BUS). |
Copied to clipboard
| Challenge: | Existing methods for abstractive summarization are unable to ensure factual consistency of generated summaries. |
| Approach: | They propose a post-editing corrector module to identify and correct factual errors in generated summaries. |
| Outcome: | The proposed model outperforms existing models on CNN/DailyMail dataset on factual consistency evaluation. |
Copied to clipboard
| Challenge: | a new method to learn which compressions to apply is based on syntactic rules for deleting spans . plausibility and salience are the two main criteria for determining which compression to apply . a recent study shows that the plausability model generally selects for grammatical and factual deletions compared to extractive methods . |
| Approach: | They propose to leave the decision about what to delete to two data-driven criteria . they show that plausibility and salience are the most important criteria if a span is deleted . |
| Outcome: | The proposed method achieves strong in-domain results on benchmark datasets and human evaluation shows that plausibility model generally selects for grammatical and factual deletions. |
Copied to clipboard
| Challenge: | Abstractive summarizations are considered to be less reliable because they distort the original meaning and can be confusing for readers. |
| Approach: | They propose a method to generate summary highlights that are understandable on their own to avoid confusion. |
| Outcome: | The proposed method allows summaries to be understood in context and avoids misdirecting readers to false conclusions. |
Copied to clipboard
| Challenge: | Abstractive text summarization has produced fluent and informative outputs, but factual inconsistency is a challenge. |
| Approach: | They propose a framework that mitigates the causal effects of language bias and irrelevancy bias by counterfactual estimation. |
| Outcome: | The proposed framework outperforms baseline methods on two widely used summarization datasets. |
Copied to clipboard
| Challenge: | Existing methods for generating abstractive summarization are inconsistent and rely on heuristically created data for error handling. |
| Approach: | They propose a contrastive learning formulation that leverages both positive and negative summaries to train summarization systems that are better at distinguishing between them. |
| Outcome: | The proposed learning framework produces more factual summaries than strong comparisons with post error correction, entailment-based reranking, and unlikelihood training. |
Copied to clipboard
| Challenge: | Experiments show that CIFLEX significantly reduces computational costs without degrading task performance. |
| Approach: | They propose a new execution system for efficient sub-task handling with a single large language model. |
| Outcome: | Experiments show that CIFLEX significantly reduces computational costs without degrading task performance. |
Copied to clipboard
| Challenge: | Scientific paper summarization is the focus of recent research . prevailing summarizing methods involve selective extraction of content from abstract, introduction, and conclusion segments within the target articles. |
| Approach: | They propose a model that incorporates references and citations to capture the impact of the document on the research community. |
| Outcome: | The proposed model generates extractive and abstractive summaries in parallel and improves their performance when considering the standard metrics. |
Copied to clipboard
| Challenge: | Existing methods for evaluating inconsistency in summarization are limited . a recent study found that more than 30% of summarized summaries are inconsistent with the source documents . |
| Approach: | They propose a method for localizing inconsistency errors in summarization using a synthetic dataset that contains factual errors likely to be produced by a common language processor. |
| Outcome: | The proposed method detects factual errors more accurately than existing weakly supervised methods . the proposed model also detects errors in original sentences more accurately . |
Copied to clipboard
| Challenge: | Political ideologies can lead people to develop misperceptions of groups with opposing opinions, such as the 2024 US presidential election, French legislative election, or the Brexit referendum. |
| Approach: | They propose a dataset and task for independently summarizing political perspectives in a set of opinionated news articles. |
| Outcome: | The proposed dataset and task evaluates models of varying sizes and architectures on a set of opinionated news articles. |
Copied to clipboard
| Challenge: | Abstractive summarization models have made great strides in recent years, but little is known about how they actually form summaries and how to understand where their decisions come from. |
| Approach: | They propose a two-step method to interpret summarization model decisions by categorizing each decoder decision into one of several generation modes. |
| Outcome: | The proposed method can identify phrases the summarization model has memorized and determine where in the training pipeline this memorization happened, and study complex generation phenomena on a per-instance basis. |
Copied to clipboard
| Challenge: | Existing abstractive summarization models do not take into account argumentative structure of legal documents, which poses a challenge towards effective abstractive summary. |
| Approach: | They propose a technique that integrates argument role labeling into the summarization process by integrating argument role labels into the document. |
| Outcome: | The proposed method improves over strong baselines with pretrained language models. |
Copied to clipboard
| Challenge: | Existing approaches to comparative reasoning rely on pretraining or fine-tuning models at the cost of massive human annotation and computation. |
| Approach: | They propose a model that prompts LLMs to generate structured intermediate comparisons by proposing aspects for comparison, followed by generating textual comparisons under each aspect. |
| Outcome: | The proposed model significantly reduces hallucination and improves consistency across various NLP tasks. |
Copied to clipboard
| Challenge: | Pretrained large language models excel in a variety of natural language processing tasks . however, they pose significant security risks due to their tendency to memorize training data . |
| Approach: | They propose a method to estimate LLM memorization using dynamic, prefix-dependent soft prompts. |
| Outcome: | The proposed method can achieve maximum relative improvement of 135.3% and 39.8% over baseline compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Preference optimization methods have been successfully applied to improve the alignment of large language models with human values. |
| Approach: | They propose to use preference optimization methods to generate rejected answers using weak LLM prompting and digit corruption to improve the mathematical reasoning abilities of language models. |
| Outcome: | The proposed method leads to increased accuracy on the GSM8K and AQuA-RAT benchmarks without annotations. |
Copied to clipboard
| Challenge: | Recent advances in efficient attention mechanisms have led to the expansion of the context length of large language models. |
| Approach: | They propose a procedure to synthesize Haystacks of documents and generate a summary that identifies relevant insights and precisely cites the source documents. |
| Outcome: | The proposed evaluation can score summaries on Coverage and Citation . the proposed evaluation lags human performance estimates by 10+ points on SummHay . |
Copied to clipboard
| Challenge: | In-context learning (ICL) is a dominant paradigm in natural language processing. |
| Approach: | They propose a prompting method for classification tasks using exemplar answers in a *comparative format' they also propose introducing a test instance before the exemplars to improve performance . |
| Outcome: | The proposed method achieves up to 13.76% increase in accuracy on classification tasks across decoder-only and encoder-decoder LLMs. |
Copied to clipboard
| Challenge: | Transformers have shown dominant performance across a range of domains including language and vision, but their computational cost grows quadratically with the sequence length, making their usage prohibitive for resource-constrained applications. |
| Approach: | They propose a segmented recurrent transformer that combines segmente recursion with recursive attention to reduce the computational cost. |
| Outcome: | The proposed model achieves higher ROUGE1 scores and lower computational complexity than current approaches. |
Copied to clipboard
| Challenge: | Large language models have demonstrated great potential in natural language generation, but their widespread adoption has raised concerns regarding content reliability and accountability. |
| Approach: | They propose a challenge to trace each sentence of a target text back to specific source sentences within potentially lengthy or multi-document inputs. |
| Outcome: | The proposed challenge traces each sentence of a target text back to specific source sentences . the dataset includes 11 scenarios covering QA and summarization in english and Chinese . |
Copied to clipboard
| Challenge: | Existing long-term open-domain dialogue datasets lack complex, real-world personalization and fail to capture implicit reasoning. |
| Approach: | They propose a large-scale long-term dataset with 2,500 examples containing approximately 100 conversation sessions to study implicit reasoning in personalized dialogues. |
| Outcome: | The proposed model improves the ability of LLMs to reason over long-term conversations with implicit contextual dependencies. |
Copied to clipboard
| Challenge: | Existing contrastive methods that ignore the context of a large language model (LLM) fail to handle instances that vary in their amount of conflict, with static methods over-adjusting when conflict is absent. |
| Approach: | They propose a fine-grained, instance-level approach called AdaCAD which dynamically adjusts the degree of conflict based on the degree. |
| Outcome: | The proposed approach outperforms baselines and improves factuality of summaries by 6.19. |
Copied to clipboard
| Challenge: | evaluating the quality of generated text is a difficult problem for large language models. |
| Approach: | They propose a dataset for multilingual, multifaceted summarization evaluation. |
| Outcome: | The proposed dataset can be used to train multilingual summarization systems . it shows that the dataset performs well on the out-of-domain meta-evaluation benchmarks TRUE and mFACE . |
Copied to clipboard
| Challenge: | Existing datasets for multi-document summarization (MDS) are either in the general domain, such as WikiSum, or very small such as DUC 1 or TAC 2011 . Existing systems for summarizing biomedical literature take 1-2 years to complete . |
| Approach: | They propose to use a multi-document summarization system based on BART to assess the quality of the summarized biomedical literature. |
| Outcome: | The proposed system has high summarization quality, but significant work remains to achieve it. |
Copied to clipboard
| Challenge: | a number of recent datasets for summarisation, scraped the web-content relying on the assumption that summary is made available with the article by the publishers. |
| Approach: | They propose a pipeline that crowd-sources summarization data and then aggressively filters the content via: automatic and partial expert evaluation. |
| Outcome: | The proposed pipeline can be applied to scraped datasets to extract better quality articles-summaries pairs. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have enabled them to process increasingly longer sequences, ranging from 2K to 2M tokens and even beyond. |
| Approach: | They propose a synthetic dataset in the financial domain that integrates Chain-of-Thought reasoning into LLMs in a supervised manner to facilitate effective long-context understanding. |
| Outcome: | The proposed model outperforms standard GPT-4o-mini on the Loong benchmark and fine tunes LLaMA-3.1-8B-Instruct on the model, achieving a 28.0% gain on the financial subset. |
Copied to clipboard
| Challenge: | Abstractive summarization systems have a severe mismatch between training and inference, i.e., exposure bias. |
| Approach: | They propose a multi-level contrastive learning framework for abstractive summarization and a tailored sparse decoder self-attention pattern to bridge the gap between training and inference. |
| Outcome: | The proposed framework outperforms the state-of-the-art models on two summarization datasets while adding relatively low overhead. |
Copied to clipboard
| Challenge: | SUM-QE is a quality estimation model for summarization that captures linguistic qualities that traditional evaluation metrics fail to capture. |
| Approach: | They propose a new quality estimation model based on BERT that addresses linguistic quality aspects that are only indirectly captured by content-based approaches to summary evaluation without comparison with human ratings. |
| Outcome: | The proposed model outperforms existing models on linguistic quality aspects that are only indirectly captured by content-based summarization evaluations without comparison with human ratings. |
Copied to clipboard
| Challenge: | Using a pretrained sequence-to-sequence language model, we explore speaker name substitution, negation scope highlighting, multi-task learning with relevant tasks, and pretraining on in-domain data. |
| Approach: | They propose a pretrained sequence-to-sequence language model that can handle different parts of dialogue belonging to multiple speakers and combine them to produce a coherent monologue summary. |
| Outcome: | The proposed techniques outperform baseline models on a dialogue summarization dataset. |
Copied to clipboard
| Challenge: | Existing methods to improve output quality without aggregating input tokens are limited by the complexity of aggregation of responses. |
| Approach: | They propose to extract and integrate segment-level commonalities from candidate samples to enhance performance of LLMs in open-ended and reasoning tasks. |
| Outcome: | The proposed method improves performance on reasoning, code generation and mathematical reasoning tasks without requiring additional models and overlooking the knowledge present among the candidates. |
Copied to clipboard
| Challenge: | Existing paradigms for multi-task training involve a shared pre-trained language model and a small, thin network (head) given an input, a target head is the head that is selected for outputting the final prediction. |
| Approach: | They examine the behaviour of non-target heads when given input that belongs to a different task than the one they were trained for. |
| Outcome: | The non-target heads exhibit emergent behaviour, which may explain the target task, or generalize beyond their original task. |
Copied to clipboard
| Challenge: | Almost all popular summarization datasets do not come with inherent quality assurance guarantees. |
| Approach: | They propose to use 5 metrics to evaluate quality of summarization datasets . they find that data usage in recent summarizing research is inconsistent with the properties of the data. |
| Outcome: | The proposed metrics can be inexpensive heuristics for detecting generically low quality examples. |
Copied to clipboard
| Challenge: | Abstractive summarization systems still include factual errors in generated summaries despite recent improvements in factuality detection . |
| Approach: | They aggregate factuality error annotations from nine existing datasets and stratify them according to the underlying summarization model. |
| Outcome: | The proposed method improves on the ChatGPT-based model and shows that it is not superior for all error types. |
Copied to clipboard
| Challenge: | Existing methods for evaluating factual consistency are primarily designed for short summaries of isolated code snippets. |
| Approach: | They propose a reference-free and fine-grained method for evaluating factual consistency in real-world code summaries. |
| Outcome: | The proposed method achieves highest correlation with human judgment among 13 baselines, improving 15-18% over the previous state-of-the-art. |
Copied to clipboard
| Challenge: | Existing MCQA datasets are small in size, which increases difficulty of model learning and generalization. |
| Approach: | They propose a multi-source meta transfer framework for low-resource multiple-choice question answering . they extend meta learning by incorporating multiple training sources to learn a generalized feature representation across domains . |
| Outcome: | The proposed framework is independent of backbone language models and can bridge the distribution gap between training sources and target. |
Copied to clipboard
| Challenge: | Existing models and datasets for training summarization models are limited for less resourceful languages like Hungarian . |
| Approach: | They propose to use a Hungarian corpus for training abstractive and extractive summarization models by cleaning, preprocessing and deduplication. |
| Outcome: | The proposed model trains abstractive and extractive summarization models using the dataset . it will be made publicly available, encouraging replication, further research, and real-world applications across various domains. |
Copied to clipboard
| Challenge: | Existing approaches to align large language models with instructions and preferences are conflicting . et al., 2023b) show that hybrid alignment training can outperform baselines . |
| Approach: | They propose a hybrid alignment training approach based on alternating alignment and modified elastic weight consolidation methods to achieve better collaboration between different alignment tasks. |
| Outcome: | The proposed approach outperforms baseline alignment training methods on summarization and dialogue tasks. |
Copied to clipboard
| Challenge: | Recent advances in large language models have improved summarization, but they still face a challenge of hallucination. |
| Approach: | They propose a taxonomy of errors to address the problem of hallucination in LLMs . they propose two prompt-based approaches for fine-grained error detection . |
| Outcome: | The proposed model outperforms existing metrics in identifying the novel "Contextual Inference" error type. |
Copied to clipboard
| Challenge: | a core part of legal work that has been underexplored in Legal NLP is the writing and editing of legal briefs. |
| Approach: | They propose to use large language models to help legal professionals with writing briefs by capturing and evaluating their abilities in language models. |
| Outcome: | The proposed tasks show that the models perform well on arguments summarization, argument completion, and case retrieval tasks. |
Copied to clipboard
| Challenge: | Debatepedia dataset limited by noise and most queries do not have relevance to document . |
| Approach: | They harness the language generation capabilities of two LLMs to regenerate queries in a Debatepedia dataset. |
| Outcome: | The proposed model can regenerate queries from the Debatepedia dataset. |
Copied to clipboard
| Challenge: | Scientific peer review is essential for the quality of academic publications. |
| Approach: | They propose a method that summarises scholarly reviews using a Rational Speech Act framework and novel uniqueness scores. |
| Outcome: | The proposed method generates more discriminative summaries than baseline methods in terms of human evaluation while achieving comparable performance with these methods in term of automatic metrics. |
Copied to clipboard
| Challenge: | Recent progress in large language models (LLMs) has revolutionized text generation. |
| Approach: | They propose a faithfulness hallucination detection model that can provide binary predictions and corresponding explanations to improve trustworthiness. |
| Outcome: | The proposed model outperforms advanced models on 12 diverse tasks. |
Copied to clipboard
| Challenge: | Large-scale generative Pre-trained Language Models (PLMs) are limited in their deployment in real-world applications. |
| Approach: | They propose to prune the feed-forward networks of generative pre-trained language models to smaller widths without designing extra operators. |
| Outcome: | The proposed method achieves 1.51x/6.96x inference speedup on GPU/CPU with 67% size reduction. |
Copied to clipboard
| Challenge: | Lack of publicly available NLG benchmarks for low-resource languages poses a challenge . authors show that IndoBART and IndoGPT achieve competitive performance on all tasks . |
| Approach: | They propose a benchmark to measure natural language generation progress in three low-resource languages of Indonesia . they use a corpus of pretraining datasets to build their models . |
| Outcome: | The proposed benchmark measures progress in Indonesian, Javanese, and Sundanese . the results highlight the importance of pretraining on closely related, localized languages . |
Copied to clipboard
| Challenge: | a recent WHO report highlights a drastic doctor-to-patient ratio . telehealth is one of the most impactful sectors where AI advances can bring a significant revolution . |
| Approach: | They propose an image-guided encoder-decoder model that uses contextual attention to create detailed visual-guides for multimodal documents. |
| Outcome: | The proposed model outperforms state-of-the-art models on multimodal question and dialogue summarization tasks. |
Copied to clipboard
| Challenge: | Existing methods to control document controllable summarization lack abundant labeled data. |
| Approach: | They propose a question-driven, unsupervised pretraining objective to improve controllability in document controllable summarization tasks. |
| Outcome: | The proposed method outperforms pre-finetuning approaches on QMSum and SQuALITY. |
Copied to clipboard
| Challenge: | Existing benchmarks measure the correlation with human judgements of faithfulness on model-generated summaries, but they are insufficient for diagnosing whether metrics are consistent, effective on human-written texts, and sensitive to different error types. |
| Approach: | They propose to use unfaithful minimal pairs to measure the consistency of automatic faithfulness metrics by comparing human-written summary pairs with a dataset of 889 human-writing, minimally different summary pairs. |
| Outcome: | The proposed benchmarks show that the most discriminative metrics tend not to be the most consistent, and that the best performing metrics are sensitive to errors. |
Copied to clipboard
| Challenge: | Large language models (LLMs) adopt autoregressive architecture, predicting the next word token based on the preceding context. |
| Approach: | They propose a method that integrates task-specific predictive models as external tools to improve model generation quality and accuracy. |
| Outcome: | The proposed method improves the generation quality and predictive accuracy of large language models in inference-driven tasks. |
Copied to clipboard
| Challenge: | Question Answering (QA) tasks require a mix of relevant and irrelevant information in these contexts to perform well. |
| Approach: | They propose a context filtering approach that removes non-essential details, summarizing crucial content through Reward Modeling. |
| Outcome: | The proposed approach outperforms baseline models in 6.8-folds. |
Copied to clipboard
| Challenge: | Existing studies have focused on the identification of social media posts that contain misrepresentations of information within associated news articles. |
| Approach: | They propose a data collection schema and curated a dataset called ManiTweet, consisting of 3.6K pairs of tweets and corresponding articles. |
| Outcome: | The proposed model outperforms large language models on the ManiTweet dataset and reveals intriguing connections between manipulation and the domain and factuality of news articles. |
Copied to clipboard
| Challenge: | Inductive transfer learning has taken the entire NLU field by storm, with models such as BERT and BART setting new state-of-the-art on countless tasks. |
| Approach: | They introduce a large-scale pretrained seq2seq model for French that is very competitive with state-of-the-art BERT-based French language models such as CamemBERT and FlauBERT. |
| Outcome: | The proposed model outperforms existing models on discriminative and generative tasks on a French summarization dataset. |
Copied to clipboard
| Challenge: | Abstractive summarization is one of the areas influenced by pre-trained language models. |
| Approach: | They propose a Transformer-based encoder-decoder model pre-trained with three novel objectives to address this issue. |
| Outcome: | The proposed model outperforms previous models on six Persian summarization tasks . it also outperformed previous models in textual entailment, question paraphrasing, and question answering . |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly used to summarize academic work . however, they can exaggerate or mischaracterize findings . |
| Approach: | They examine how Narrative License (NL) emerges in large language models summaries . authors find that stated stances and user personas produce predictable shifts . |
| Outcome: | The proposed models can exaggerate or mischaracterize findings in scholarly articles . the authors show that the models' "sycophancy" can reduce NL . |
Copied to clipboard
| Challenge: | Automated evaluation metrics are an essential part of the development of text-generation tasks such as summarization. |
| Approach: | They propose to use top-scoring system outputs to assess the reliability of automatic evaluation metrics for text summarization. |
| Outcome: | The proposed evaluation method is based on human judgments from 25 top-scoring neural summarization systems. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown remarkable capabilities in tasks such as summarization, arithmetic reasoning, and question answering. |
| Approach: | They propose a framework that explores decisions’ consequences from multiple stakeholder perspectives and a SKIG framework to enhance moral reasoning in large language models. |
| Outcome: | The proposed framework exhibits marked improvements compared to baselines across different language models and benchmarks. |
Copied to clipboard
| Challenge: | Existing systems that can recognize spoken content, extract key information, and produce concise summaries are lacking in meeting transcription and summarization. |
| Approach: | They propose a multimodal dataset that integrates information from speech, vision, and text modalities to facilitate automatic meeting transcription and summarization (AMTS). |
| Outcome: | The proposed dataset reduces the character error rate (CER) by 36.60% to 20.27% and improves speech recognition and large language models. |
Copied to clipboard
| Challenge: | Comparative reasoning is a process of comparing objects, concepts, or entities to draw conclusions. |
| Approach: | They propose a framework to pre-train language models for enhancing comparative reasoning abilities . they collect scalable data for text-based entity comparison . |
| Outcome: | The proposed framework significantly improves comparative reasoning abilities under low-resource conditions on downstream tasks. |
Copied to clipboard
| Challenge: | Abstractive summarization models (LLMs) have demonstrated impressive performance in various tasks, but they are still suffering from factual inconsistency problem called hallucination. |
| Approach: | They propose to improve the faithfulness of large language models by impelling them to process the entire article more fairly and faithfully. |
| Outcome: | The proposed strategy improves the faithfulness of large language models in summarization while maintaining their fluency and informativeness. |
Copied to clipboard
| Challenge: | Existing multilingual Large Language Models are not specifically trained with objectives for managing code-switching scenarios. |
| Approach: | They propose to use multilingual Large Language Models to perform sentiment analysis, machine translation, summarization and word-level language identification to compare their performance to fine-tuned models of much smaller scales. |
| Outcome: | The proposed models show that they underperform in comparison to fine-tuned models of much smaller scales. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for monolingual summarization require translation to evaluate the factuality of cross-lingual summmarization. |
| Approach: | They propose to analyze cross-lingual factuality by collecting annotations and generated summaries from models at summary level and sentence level. |
| Outcome: | The proposed dataset shows that over 50% of generated summaries contain factual errors with different characteristics from monolingual summarization. |
Copied to clipboard
| Challenge: | Existing methods for dialogue summarization are far from satisfactory . omission is a major factor in affecting the quality of summarizing, but few studies have explored the problem . |
| Approach: | They propose a dataset that provides high-quality omission labels for dialogue summarization . they propose to use this dataset to detect omitted dialogue utterances . |
| Outcome: | The proposed dataset improves summarization quality by providing ground-truth omission labels . the proposed dataset and codes are publicly available . |
Copied to clipboard
| Challenge: | Existing models do not capture factors that contribute to producing consistent text. |
| Approach: | They propose a benchmark test to evaluate text complexity in generative models by observing linguistic properties of input prompts. |
| Outcome: | The proposed model fails to preserve complexity of input prompts even if finetuned with professionally written texts. |
Copied to clipboard
| Challenge: | Existing solutions for word probability distributions are limited and the output softmax layer is inherently limited. |
| Approach: | They propose to use the output softmax layer to compute the word probability distribution instead of using pointer networks to break the bottleneck. |
| Outcome: | The proposed method improves factCC score by 2 points in CNN/DM and XSUM dataset, and MAUVE scores by 30% in bookSum paragraph-level dataset. |
Copied to clipboard
| Challenge: | Existing approaches to meeting summarization are limited due to noise, lengthy transcripts, and scattered salient information. |
| Approach: | They propose a two-step framework for meeting summarization that leverages a self-supervised paradigm to reconstruct transcripts and a relative positional bucketing algorithm to equip models to generate the summary. |
| Outcome: | The proposed method significantly reduces memory consumption and processing time on two meeting summarization datasets. |
Copied to clipboard
| Challenge: | Current abstractive summarization systems tend to hallucinate unfaithful content . however, the most common method does not disentangle factual errors from other errors. |
| Approach: | They propose a back-translation-style approach to augment negative samples that mimic factual errors made by the model. |
| Outcome: | The proposed method improves faithfulness without sacrificing informativeness . it incorporates negative samples into training, and produces faithful/unfaithful summaries . |
Copied to clipboard
| Challenge: | Summarizing text is not a straightforward task. |
| Approach: | They propose to use automated transcriptions to generate reports from automatic transcriptions as a dataset for neural summarization. |
| Outcome: | The proposed model improves on publicmeetings corpus on a dataset of aligned public meetings. |
Copied to clipboard
| Challenge: | Meeting transcripts are a promising domain for natural language tasks . lack of annotated data impedes research on other important tasks in this domain . |
| Approach: | They propose an extractive QA dataset comprising questions asked by meeting participants and corresponding responses. |
| Outcome: | The proposed dataset extracts questions asked by meeting participants and corresponding responses from transcripts. |
Copied to clipboard
| Challenge: | Query-focused summarization (QFS) is gaining prominence in research community. |
| Approach: | They propose to integrate Learning-to-Rank (LTR) with Query-focused Summarization (QFS) to enhance the summary relevance via content prioritization. |
| Outcome: | The proposed model outperforms the state-of-the-art on QMSum benchmark and SQuALITY benchmark while offering a lower training overhead. |
Copied to clipboard
| Challenge: | Controllable summarization is a form of outputs that tailors summaries to user-specified attributes. |
| Approach: | They propose an adaptive planning framework that reframes the task as planning the order of sequential attribute control with a customized Monte Carlo Tree Search. |
| Outcome: | The proposed framework surpasses LLM-based self-planning models and fine-tuned baselines in multi-attribute controllable summarization. |
Copied to clipboard
| Challenge: | Existing studies have focused on specialized BERT-variants and recent LLMs to reason inconsistencies. |
| Approach: | They propose to incorporate task-specific taxonomy into inferences to facilitate both zero-shot and supervised paradigms. |
| Outcome: | The proposed model outperforms specialized non-LLM and recent LLM models in a number of domains. |
Copied to clipboard
| Challenge: | Uncertainty quantification (UQ) provides measures of uncertainty, such as an estimate of the confidence in an LLM’s generated output. |
| Approach: | They propose a black-box approach where consistency is used as a proxy for confidence in a model's output. |
| Outcome: | The proposed methods are primarily but not necessarily entirely black- box, with consistency between output and other sampled generations used as a proxy for confidence in its correctness. |
Copied to clipboard
| Challenge: | Summarization evaluation approaches have relied on ROUGE for summarization, but they fall short of human evaluations. |
| Approach: | They propose a new approach to evaluate summaries by leveraging retrieval techniques . they use a dual-encoder retrieval setup to train a retrieval task . |
| Outcome: | The proposed method outperforms existing methods on two document summarization benchmarks and a long document summmarization test. |
Copied to clipboard
| Challenge: | Current work on understanding assembly code is oriented towards generating function names, which involve numerous abbreviations that make them confusing. |
| Approach: | They propose a control flow graph and pseudo code guided binary code summarization framework to learn the comprehensive binary function execution behavior and logic semantics. |
| Outcome: | The proposed framework improves the efficiency of reverse engineering on 3 different binary optimization levels for 3 different computer architectures. |
Copied to clipboard
| Challenge: | In-Context Learning (ICL) is a key method in prompt engineering, but its long retrieved contexts and limited token throughput will slow reasoning speeds. |
| Approach: | They propose a method that leverages the overlap between context and model output to generate drafts from the context. |
| Outcome: | The proposed method achieves the highest mean speedup on Vicuna-7B, Llama2-7B-Chat, and Llma3-8B-Instruct tasks. |
Copied to clipboard
| Challenge: | Summarization of poetry is a challenging task as it can be easily lost if only the literal meaning is considered. |
| Approach: | They propose to use poetry as a model to summarize poetry and provide a dataset to evaluate their creative language interpretation capacity. |
| Outcome: | The proposed dataset consisting of 3011 samples and its corresponding summarized interpretation in the English language provides an opportunity to evaluate the creative language interpretation capacity of the proposed models. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have revolutionized natural language processing with impressive capabilities, but they lack domain specificity, real-time information and face challenges in solving specialized problems. |
| Approach: | They propose a multi-LLM approach that decomposes the aforementioned capabilities into a planner, caller, and summarizer. |
| Outcome: | The proposed model outperforms existing models by demonstrating its effectiveness and advantages in tool learning. |
Copied to clipboard
| Challenge: | n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear. |
| Approach: | They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics. |
| Outcome: | The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand. |
Copied to clipboard
| Challenge: | Existing methods for adapting LLMs to low-resource tasks keep LoRA parameters frozen and the low-level problem out of their scope. |
| Approach: | They propose a LoRA merge method that updates and prunes LoRA parameters through fine-tuning with minimal target task data. |
| Outcome: | The proposed method improves performance on a low-resource language generation task and improves on previous methods. |
Copied to clipboard
| Challenge: | Large language models have revolutionized the field of NLP by achieving state-of-the-art performance on various tasks. |
| Approach: | They investigate the membership inference attack by using model's API to determine if a sample was part of the training data. |
| Outcome: | The proposed model is able to identify if a sample was part of the training data and exploits its similarity and resistance to document modifications as potential MI signals on widely used datasets. |
Copied to clipboard
| Challenge: | Existing models for multilingual generation lack thorough analysis due to extensive linguistic diversity. |
| Approach: | They propose to classify multilingual generation methodologies into three categories based on their underlying modeling principles . they introduce an automatic metric to mitigate spurious correlations associated with language mixing . |
| Outcome: | The proposed model improves in high-resource, low-resourced, and zero-shot scenarios. |
Copied to clipboard
| Challenge: | Product review summarization aims to generate a concise summary based on product reviews . factual accuracy, aspect comprehensiveness, and content relevance are challenges . |
| Approach: | They propose an FB-Thinker framework to improve product review summarization ability . they propose two Chinese product review summary datasets for instruction-tuning and evaluation . |
| Outcome: | The proposed framework improves product review summarization with forward reasoning and backward refinement. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are crucial for enabling intelligent experiences across applications. |
| Approach: | They propose a low-rank adaptive localization method that uses rank-norm regularization to determine the optimal rank for each weight matrix. |
| Outcome: | NormAL LoRA reduces adapter parameters by 37% while preserving full fine-tuning performance. |
Copied to clipboard
| Challenge: | Existing studies have shown that automated evaluation approaches correlate with human ratings in English, but this is unclear for other languages. |
| Approach: | They construct a small-scale pilot dataset containing article-summary pairs and human ratings in English, Chinese and Indonesian to measure the strength of summaries. |
| Outcome: | The results show that standard metrics are unreliable measures of quality in Chinese and Indonesian. |
Copied to clipboard
| Challenge: | Using textual feedback, language models can be trained to learn from textual inputs. |
| Approach: | They propose an approach that aligns language models with user preferences expressed in text. |
| Outcome: | The proposed approach outperforms PPO on toxicity reduction, summarization, and dialog response tasks while achieving the same performance with only 20% of the samples. |
Copied to clipboard
| Challenge: | Existing studies have focused on summarizing factual information, leaving out affective content. |
| Approach: | They propose to quantify the preservation of affective content in dialogue summaries using PSentScore measures. |
| Outcome: | The proposed measures show that state-of-the-art summarization models do not preserve well affective content in their summaries. |
Copied to clipboard
| Challenge: | a recent study has found that disfluencies negatively impact spoken content summarization . |
| Approach: | They aim to quantify the impact of disfluency on spoken content summarization . they also investigate two methods towards improving summarizing in the presence of disflouencies . |
| Outcome: | The proposed methods improve summarization quality in the presence of disfluencies. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have become increasingly prevalent in the field of Natural Language Processing (NLP), achieving unprecedented performance across linguistic tasks. |
| Approach: | They propose a framework to quantify and analyze context-driven over-refusal . they find that over-fusals depend on the task, system prompts, model family, and the number of retrieved documents. |
| Outcome: | The proposed framework quantifyes and analyzes the concept of context-driven over-refusal on two public corpora. |
Copied to clipboard
| Challenge: | MoE-based LLMs are not explicitly supervised to select suitable experts. |
| Approach: | They propose Exploration-Driven Reinforcement Learning (ERL) which explicitly optimizes the router by exploration of alternative routing paths. |
| Outcome: | The proposed method improves summarization (SAMSum, XSUM, question answering, and language modeling), and raises routing quality, delivering 8.9 higher MRR than baselines over 100 perturbed routing paths. |
Copied to clipboard
| Challenge: | a recent study has focused on detecting media bias in news articles . a multi-document event relation graph is used to generate a neutralized summary . |
| Approach: | They propose to generate a neutralized summary given multiple articles presenting different ideological views. |
| Outcome: | The proposed method mitigates media bias and improves content preservation. |
Copied to clipboard
| Challenge: | Existing task-aware methods require loading the entire input sequence at once for compression, which suffer from computational inefficiency. |
| Approach: | They propose a framework that adopts an adaptive hybrid reading strategy to reduce computational inefficiency and redundant information in long-context scenarios. |
| Outcome: | Experiments show that RAM outperforms baselines on multiple question answering and summarization benchmarks while delivering up to a 12x speedup on long inputs. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have achieved remarkable performance across NLP tasks . however, in long-context scenarios, they face high computational cost and information redundancy. |
| Approach: | They propose an encoder-decoder context compression framework that generates a compact sequence of soft tokens for downstream tasks. |
| Outcome: | Experiments show that GMSA outperforms baselines on multiple long-context question answering and summarization benchmarks while maintaining low end-to-end latency. |
Copied to clipboard
| Challenge: | Podcasts and other audiovisual content are becoming more and more a part of everyday communication and the digital age is changing from text to voice. |
| Approach: | They synthesize the current state of the field and highlight the need for realistic evaluation benchmarks and multilingual datasets. |
| Outcome: | The proposed frameworks are based on evaluation protocols and datasets and highlight the need for realistic benchmarks and multilingual datasets. |
Copied to clipboard
| Challenge: | Large language models exhibit the _”lost in the middle” phenomenon when they are unevenly attending to different parts of the provided context. |
| Approach: | They propose principled content selection as a way to increase source coverage . they use determinantal point processes to prioritize diverse content . |
| Outcome: | The proposed method improves source coverage on the DiverseSumm benchmark. |
Copied to clipboard
| Challenge: | despite LLMs becoming increasingly multilingual, most studies on detecting and quantifying LLM hallucination are English-centric . |
| Approach: | They train a multilingual hallucination detection model and conduct a large-scale study across 30 languages and 6 open-source LLM families. |
| Outcome: | The proposed model is based on an English-centric model and annotates gold data for five high-resource languages. |
Copied to clipboard
| Challenge: | Effective protection of private information is essential for knowledge dissemination in sensitive domains such as medical and legal. |
| Approach: | They perform a comprehensive study of privacy risks in LM-based summarization using closed- and four-weight models of different sizes and families. |
| Outcome: | The proposed models show that they leak personally identifiable information in their summaries, compared to human-generated summary summators, which show significantly higher privacy protection levels. |
Copied to clipboard
| Challenge: | Existing approaches to personalize large language models (LLMs) rely on heuristic methods to compress user profiles but they ignore how LLMs process and prioritize different profile components. |
| Approach: | They propose an attention-guided context compression framework that leverages attention feedback from a marking model to mark important personalization sentences and guides a compression model to generate task-relevant compressed user contexts. |
| Outcome: | The proposed framework outperforms baselines across tasks, token limits, and settings while reducing token usage by 50 times. |
Copied to clipboard
| Challenge: | Text revision is a core process in document creation, capturing how authors iteratively refine, reorganize, and improve written content. |
| Approach: | They synthesize text revision research through the lens of edit intentions . they review prior work across the revision workflow including corpus construction, edit intention taxonomies, edit intentions, and edit intention identification. |
| Outcome: | The proposed approach synthesizes datasets, taxonomies, identification methods, and applications and highlights key open research directions. |
Copied to clipboard
| Challenge: | Recent advances in summary evaluation are based on model-based metrics to assess quality dimensions, such as completeness, conciseness, and faithfulness. |
| Approach: | They propose a general framework that generates individual and average proxy scores without relying on reference summaries, human annotations, or expensive model-based metrics. |
| Outcome: | The proposed framework outperforms baselines on seven datasets on continuous-value scenarios, such as summarization, but is applicable to discrete-value tasks, such QA. |
Copied to clipboard
| Challenge: | Traditional metrics like BLEU and BERTScore fail to capture semantic fidelity in generative text-to-text tasks. |
| Approach: | They propose a cross-examination framework that generates verifiable questions from each text and performs a Cross-exam to derive three interpretable scores: Coverage, Conformity, and Consistency. |
| Outcome: | The proposed framework detects critical errors across translation, summarization and clinical note-generation and human expert validation shows it is reliable without gold references. |
Copied to clipboard
| Challenge: | Recent advances in summarization focus on improving summary quality across multiple dimensions, but they overlook the challenge of controlling summary generation with respect to individual dimensions. |
| Approach: | They propose a loss function that aligns model outputs with fine-grained, model-based evaluation scores to enable both improvement in summary quality and dimension-specific control. |
| Outcome: | The proposed method improves the overall quality of summaries while maintaining strong control over individual quality dimensions. |
Copied to clipboard
| Challenge: | Recent advances in autonomous digital agents highlight their potential for structured tasks through autonomous decision-making and task decomposition, but it remains unclear how well such systems support real-world information-intensive workflows. |
| Approach: | They propose a benchmark to evaluate how journalists can use agents to organize and organize information from the web. |
| Outcome: | The proposed system can be used to iterate and evaluate newswriting tasks in real-world situations. |
Copied to clipboard
| Challenge: | a new approach to adapt generalist models to expert domains is needed to overcome this problem. |
| Approach: | They propose a parameter-efficient domain adaptation approach that combines vocabulary adaptation with pretraining for LLM-based text summarization. |
| Outcome: | The proposed approach reduces training time by 35-55% over continual pretraining and reduces parameter counts up to 37% w.r.t expansion-only methods. |
Copied to clipboard
| Challenge: | Existing agentic applications rely on LLMs to self-assess the factuality of outputs . but current LLM systems fail to detect hallucinations . |
| Approach: | They propose a benchmark that breaks down hallucination detection into four critical steps . they show that when halluciation detection is treated as a multi-step process, all models achieve considerably better performance. |
| Outcome: | The proposed benchmark breaks down hallucination detection into four critical steps . it shows that when halluciation detection is treated as a multi-step process, all models achieve considerably better performance. |
Copied to clipboard
| Challenge: | Large language models produce content that contradicts or overlooks information provided in the input context, a phenomenon known as faithfulness hallucination. |
| Approach: | They propose a lightweight framework that boosts the generation probability of context-relevant tokens by boosting the generation of tokens. |
| Outcome: | The proposed framework improves faithfulness metrics with minimal generation overhead. |