Challenge: Prior approaches for unsupervised keyphrase extraction relied on heuristic notions of phrase importance via embedding clustering or graph centrality.
Approach: They propose an approach which defines keyphrases as document phrases that are salient for predicting the topic of the document.
Outcome: The proposed method alleviates the need for ad-hoc heuristics and achieves state-of-the-art results in scientific publications and news articles.

Similar Papers

Unsupervised Keyphrase Extraction by Learning Neural Keyphrase Set Function (2023.findings-acl)

Copied to clipboard

Challenge: Unsupervised keyphrase extraction is a task of extracting a keyphrase set that provides readers with highlevel information about the key ideas or important topics described in the document.
Approach: They propose an unsupervised keyphrase extraction task that is a document-set matching problem instead of modeling the relevance between an individual phrase and the document.
Outcome: The proposed model outperforms the state-of-the-art unsupervised keyphrase extraction baselines by a large margin.
A Survey on Recent Advances in Keyphrase Extraction from Pre-trained Language Models (2023.findings-eacl)

Copied to clipboard

Challenge: Keyphrase extraction is a key component in Natural Language Processing (NLP) systems for selecting a set of phrases from the document that could summarize the important information discussed in the source document.
Approach: They propose to use supervised and unsupervised keyphrase extraction techniques to investigate the state-of-the-art models for keyphrase extracting.
Outcome: The proposed keyphrase extraction system can significantly accelerate the speed of retrieval and help people get first-hand information from a long document quickly and accurately.
Improving Embedding-based Unsupervised Keyphrase Extraction by Incorporating Structural Information (2023.findings-acl)

Copied to clipboard

Challenge: Existing unsupervised keyphrase extraction models ignore the indicative role of the highlights in certain locations, leading to wrong keyphrases extraction.
Approach: They propose a Highlight-Guided Unsupervised Keyphrase Extraction model that models phrase-document relevance via the highlights of documents and calculates cross-phrase relevance between all candidate phrases.
Outcome: The proposed model outperforms the state-of-the-art unsupervised keyphrase extraction models on three benchmarks.
AttentionRank: Unsupervised Keyphrase Extraction using Self and Cross Attentions (2021.emnlp-main)

Copied to clipboard

Challenge: Keyword or keyphrase extraction is to identify words or phrases presenting the main topics of a document.
Approach: They propose a hybrid attention model to identify keyphrases from a document in an unsupervised manner.
Outcome: The proposed model is effective and robust on long and short documents.
Unsupervised Keyphrase Extraction with Multipartite Graphs (N18-2)

Copied to clipboard

Challenge: Recent years have witnessed a resurgence of interest in automatic keyphrase extraction.
Approach: They propose an unsupervised keyphrase extraction model that encodes topical information within a multipartite graph structure.
Outcome: The proposed model improves on three widely used datasets.
Key2Vec: Automatic Ranked Keyphrase Extraction from Scientific Articles using Phrase Embeddings (N18-2)

Copied to clipboard

Challenge: Keyphrase extraction is a fundamental task in natural language processing that facilitates mapping of documents to a set of representative phrases.
Approach: They propose an unsupervised technique that leverages phrase embeddings for ranking keyphrases extracted from scientific articles using theme-weighted PageRank.
Outcome: The proposed method performs better on benchmark datasets than other methods and is of high quality.
Attention-Seeker: Dynamic Self-Attention Scoring for Unsupervised Keyphrase Extraction (2025.coling-main)

Copied to clipboard

Challenge: Unsupervised keyphrase extraction methods require large amounts of labeled data and are often domainspecific, limiting their practical applicability.
Approach: They propose an unsupervised keyphrase extraction method that leverages self-attention maps from a Large Language Model to estimate the importance of candidate phrases.
Outcome: The proposed method outperforms baseline models on four datasets and is highly efficient on three of four dataset.
SAMRank: Unsupervised Keyphrase Extraction using Self-Attention Map in BERT and GPT-2 (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for keyphrase extraction use contextualized embeddings to capture semantic relevance between words, sentences, and documents.
Approach: They propose an unsupervised keyphrase extraction approach that uses only a self-attention map in a pre-trained language model to determine the importance of phrases.
Outcome: The proposed approach outperforms embedding-based models on three keyphrase extraction datasets.
Mitigating Over-Generation for Unsupervised Keyphrase Extraction with Heterogeneous Centrality Detection (2023.emnlp-main)

Copied to clipboard

Challenge: Existing keyphrase extraction models incorrectly determine a keyphrase as a phrase but output other candidates as keyphrases because they contain the same word.
Approach: They propose a new approach that detects both implicit and explicit centrality within a heterogeneous graph as the importance score of each candidate keyphrase.
Outcome: The proposed approach outperforms state-of-the-art keyphrase extraction models on three benchmark datasets.
Exploiting Position and Contextual Word Embeddings for Keyphrase Extraction from Scientific Papers (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for keyphrase extraction are either supervised or unsupervised.
Approach: They propose an unsupervised algorithm that exploits contextual word embeddings and positional information to create a biased PageRank.
Outcome: The proposed algorithm outperforms previous approaches and strong baselines on five benchmark datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations