Term Set Expansion based NLP Architect by Intel AI Lab (D18-2)

Copied to clipboard

Challenge: SetExpander is a corpus-based system for expanding a seed set of terms into a more complete set of words belonging to the same semantic class.
Approach: They propose a corpus-based system for expanding a seed set of terms into a more complete set of words that belong to the same semantic class.
Outcome: The proposed system can expand a seed set of terms into a more complete set of words belonging to the same semantic class.

Similar Papers

SetExpander: End-to-end Term Set Expansion Based on Multi-Context Term Embeddings (C18-2)

Copied to clipboard

Challenge: SetExpander is a corpus-based system for expanding a seed set of terms into a more complete set of words belonging to the same semantic class.
Approach: They propose to use a corpus-based system for expanding a seed set of terms into a more complete set of words that belong to the same semantic class.
Outcome: The proposed system can expand a seed set of terms, validate it, re-expand the expanded set and store it, thus simplifying the extraction of domain-specific fine-grained semantic classes.
A Two-Stage Masked LM Method for Term Set Expansion (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for Term Set Expansion are either distributional or pattern-based . Term set expansion is a task of expanding a small seed set of example terms into a larger set of terms that belong to the same semantic category.
Approach: They propose a method which uses neural masked language models to expand a small seed set of terms into a larger set of semantic terms.
Outcome: The proposed method outperforms state-of-the-art methods due to the small seed set size . it uses neural masked language models to query large, pre-trained mlms .
Empower Entity Set Expansion via Language Model Probing (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for expanding seed entities with new entities belong to the same semantic class are difficult to implement and can lead to accumulative errors.
Approach: They propose an iterative set expansion framework that leverages automatically generated class names to address the semantic drift issue.
Outcome: The proposed framework generates high-quality class names and outperforms state-of-the-art methods significantly.
Distributional Term Set Expansion (L18-1)

Copied to clipboard

Challenge: Iterative term set expansion methods for distributional semantic models are used to label terms belonging to a sought after term set.
Approach: They compare iterative term set expansion methods for distributional semantic models to the Simple Margin method, an active learning approach to classification using Support Vector Machines.
Outcome: The proposed methods outperform centrality and classification based methods for distributional semantic models over five different term sets.
SynSetExpan: An Iterative Framework for Joint Entity Set Expansion and Synonym Discovery (2020.emnlp-main)

Copied to clipboard

Challenge: Entity set expansion and synonym discovery are two critical NLP tasks that are often performed separately, without exploring their interdependencies.
Approach: They propose a framework that enables two tasks to mutually enhance each other by including popular entities’ infrequent synonyms into the set, which boosts set expansion recall.
Outcome: The proposed framework can be used to enhance two NLP tasks by including popular entities’ infrequent synonyms into the set, which boosts set expansion recall.
A Unified Taxonomy-Guided Instruction Tuning Framework for Entity Set Expansion and Taxonomy Expansion (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies view entity set expansion, taxonomy expansion, and seed-guided taxonomies as three separate tasks.
Approach: They propose a taxonomy-guided instruction tuning framework to teach a large language model to generate siblings and parents for query entities.
Outcome: The proposed framework outperforms baselines on multiple benchmark datasets.
Topic Taxonomy Expansion via Hierarchy-Aware Topic Phrase Generation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for topic taxonomies focus on frequent terms and local topic-subtopic relations, which leads to limited topic term coverage.
Approach: They propose a framework for topic taxonomy expansion that directly generates topic-related terms belonging to new topics.
Outcome: The proposed framework outperforms baseline methods on two real-world text corpora.
Low-resource Entity Set Expansion: A Comprehensive Study on User-generated Text (2022.findings-naacl)

Copied to clipboard

Challenge: Existing benchmarks for entity set expansion (ESE) are limited to well-formed text and well-defined concepts.
Approach: They propose to use user-generated text to assess the generalizability of ESE methods by identifying phenomena such as non-named entities, multifaceted entities and vague concepts.
Outcome: The proposed methods are based on user-generated text to assess their generalizability and performance.
Bringing Emerging Architectures to Sequence Labeling in NLP (2026.eacl-long)

Copied to clipboard

Challenge: Pretrained Transformer encoders are the dominant approach to sequence labeling . however, few have been applied to sequence labels on flat or simplified tasks .
Approach: They propose to use pretrained Transformer encoders to model relations across words . they find that the architectures adapt well across tagging tasks that vary in complexity .
Outcome: The proposed architectures perform well across tagging tasks across languages and datasets.
Transforming Term Extraction: Transformer-Based Approaches to Multilingual Term Extraction Across Domains (2021.findings-acl)

Copied to clipboard

Challenge: Automated Term Extraction (ATE) is a challenging task, with few exceptions.
Approach: They propose to use a transformer-based term extraction model to extract terms from sentences . they also propose to employ a language model for token classification and a sequence model to reduce sentences to terms .
Outcome: The proposed models outperform baselines on the ATE challenge TermEval 2020 dataset in English, French, and Dutch.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations