Challenge: Existing knowledge repositories rely on Wikipedia for core sets of topics and knowledge assertions.
Approach: They propose an open-domain method for automatically annotating modifier constituents (20th-century’) within Wikipedia categories with properties (date of birth).
Outcome: The proposed method improves precision and recall over a set of Wikipedia categories.

Similar Papers

A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.
SWiPE: A Dataset for Document-Level Simplification of Wikipedia Pages (2023.acl-long)

Copied to clipboard

Challenge: Prior work on document-level simplification has focused on sentence-level edits, while many desirable edits require document- level context.
Approach: They propose a dataset that reconstructs the document-level editing process from English Wikipedia to paired Simple Wikipedia articles.
Outcome: The proposed dataset reconstructs the document-level editing process from English Wikipedia (EW) articles to paired Simple Wikipedia (SEW) pages.
Contribution of Move Structure to Automatic Genre Identification: An Annotated Corpus of French Tourism Websites (2024.lrec-main)

Copied to clipboard

Challenge: a concept of move structure has been overlooked in genre analysis, but it is not widely used in natural language processing.
Approach: They propose to incorporate move structure into a neural architecture for automatic genre identification.
Outcome: The proposed approach can increase performance and reduce computational power.
Possessors Change Over Time: A Case Study with Artworks (D18-1)

Copied to clipboard

Challenge: Existing methods to extract possession relations from Wikipedia articles can be used to extract possessors over time.
Approach: They propose to extract possession relations from Wikipedia articles and temporal information indicating when these relations are true.
Outcome: The proposed annotation scheme yields many possessors over time for a given artwork, and an LSTM ensemble can automate the task.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
Automatic Annotation of Semantic Term Types in the Complete ACL Anthology Reference Corpus (L18-1)

Copied to clipboard

Challenge: a recent increase in quantitative studies of scientific text collections has led to a significant increase in the use of semantic labeling techniques.
Approach: They propose to use semantic class labels to enhance a well-known resource . they use semantic labels to assign semantic class labeling to technical terms .
Outcome: The proposed approach enhances the ACL Anthology Reference Corpus with semantic class labels for 20,000 technical terms . the goal is to use this information as one feature in the profiling of scientific papers, communities, and disciplines.
ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links (2026.eacl-long)

Copied to clipboard

Challenge: Using retrieval models and LLMs achieves a 73% approval rate for suggested links, more than doubling the acceptance of strong retrievers alone.
Approach: They propose a domain-agnostic framework for bootstrapping sentence-level cross-document links from scratch and apply it to large-scale human-in-the-loop annotation of natural text pairs.
Outcome: The proposed framework generates semi-synthetic datasets and uses them to benchmark and shortlist the best-performing methods and applies them in large-scale human-in-the-loop annotation of natural text pairs.
Semantic Supersenses for English Possessives (L18-1)

Copied to clipboard

Challenge: Existing semantic categories for possessive constructions are limited to nominals and s-genitives.
Approach: They propose to use a supersense inventory to annotate English possessives . they show existing supersensor categories are readily applicable to possessives.
Outcome: The proposed annotations are applied to English possessives in a corpus of web reviews.
Explicit Semantic Decomposition for Definition Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing definition generation methods rely on decoding to extract semantic components of words.
Approach: They propose a method which explicitly decomposes meaning of words into semantic components and models them with discrete latent variables for definition generation.
Outcome: The proposed method outperforms existing methods on WordNet and Oxford benchmarks.
Is a Document Educational or Just Wikipedia-Style? — Pitfalls of Classifier-Based Quality Filtering (2026.acl-short)

Copied to clipboard

Challenge: Large Language Models (LLMs) are pre-trained on massive data corpora, and the quality of these corporales is one of the main factors in achieving stateof-the-art performance.
Approach: They propose to use Wikipedia-style reformatting to alter a model's quality assessment and enable low-quality content to surpass filtering thresholds.
Outcome: The proposed model would reverse filtering decision for approximately 7% of evaluated documents, thereby admitting content into the pre-training corpus that would otherwise have been excluded.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations