Corpus-based Content Construction (C18-1)

Copied to clipboard

Challenge: Existing work in this direction focuses on generating content for standard platforms like Wikipedia, where the content style is fairly consistent, but there could be multiple representations of the same information across the repository.
Approach: They propose an automatic approach to generate an initial version of the author’s intended text based on an input content snippet.
Outcome: The proposed approach improves performance against baselines on several metrics.

Similar Papers

Effective and Efficient Query-aware Snippet Extraction for Web Search (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods to extract webpage snippets ignore contextual information of webpages, which may be sub-optimal.
Approach: They propose a query-aware webpage snippet extraction method called DeepQSE that captures contextual information of webpages.
Outcome: The proposed method can significantly improve the performance of DeepQSE without affecting its performance.
Principled Content Selection to Generate Diverse and Personalized Multi-Document Summaries (2025.acl-long)

Copied to clipboard

Challenge: Large language models exhibit the _”lost in the middle” phenomenon when they are unevenly attending to different parts of the provided context.
Approach: They propose principled content selection as a way to increase source coverage . they use determinantal point processes to prioritize diverse content .
Outcome: The proposed method improves source coverage on the DiverseSumm benchmark.
Automatic Document Sketching: Generating Drafts from Analogous Texts (2021.findings-acl)

Copied to clipboard

Challenge: Large pre-trained language models have made it possible to make high-quality predictions on how to add or change a sentence in a document.
Approach: They propose a task to generate entire draft documents for the writer to review and revise.
Outcome: The proposed model can make high-quality predictions on how to add or change a sentence in a document, but it lacks the branching factor to offer useful editing suggestions at a global or document level.
Neural Models for Documents with Metadata (P18-1)

Copied to clipboard

Challenge: specialized models are often used to model text corpora without metadata . specialized algorithms are not widely used in the digital humanities and political science fields .
Approach: They propose a general neural framework based on topic models to enable customization of metadata.
Outcome: The proposed framework achieves strong performance with a manageable tradeoff between perplexity, coherence, and sparsity.
AMR Beyond the Sentence: the Multi-sentence AMR corpus (C18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is limited to capturing the semantics of individual sentences.
Approach: They propose a corpus that annotates coreference and similar phenomena on top of existing AMRs.
Outcome: The proposed corpus is compared with existing corpora on sentence-level semantics . it shows that it can be used for information extraction and question answering .
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Text Mining for History: first steps on building a large dataset (L18-1)

Copied to clipboard

Challenge: a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way .
Approach: They propose to use a Brazilian historical-biographical dictionary as a resource for text mining.
Outcome: The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated .
Summarization Beyond News: The Automatically Acquired Fandom Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Abstractive summarization methods require large corpora to train neural architectures.
Approach: They propose a novel automatic corpus construction approach that automatically constructs large open-licensed summarization corpora from existing large text collections and an evaluation process with human annotators.
Outcome: The proposed approach can be used to train abstractive summarization models on large corpora and through a manual evaluation with human annotators.
Text-to-Text Automatic Story Generation: A Survey (2026.eacl-srw)

Copied to clipboard

Challenge: Automated story generation aims to produce coherent, engaging, and contextually consistent narratives with minimal or no human involvement . despite advances in large language models, maintaining narrative coherence, character consistency, storyline diversity, and plot controllability in generating stories is still challenging.
Approach: They propose to develop new evaluation metrics and better data sets to support automatic story generation.
Outcome: The proposed evaluation metrics and better datasets will improve narrative coherence and consistency and explore practical applications of story generation.
Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing (2026.acl-long)

Copied to clipboard

Challenge: Annotated corpora are crucial in the field of natural language processing, but are difficult to exchange among researchers.
Approach: They propose a method to lawfully share the annotations of any sequential copyrighted corpus.
Outcome: The proposed method is robust to reasonable divergences in the version of the copyrighted data owned by the user.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations