Challenge: Detecting novelty of an entire document is an AI frontier problem . present state-of-the-art text matching techniques are unable to process such redundancy.
Approach: They propose a document-level novelty detection resource that can be used to benchmark techniques . they crawl news documents across several domains and use it to find out whether they contain new information .
Outcome: The proposed dataset is compared with a standard system for document novelty detection . the proposed system can detect elements that have not appeared before, or new or original .

Similar Papers

Semantic Novelty Detection and Characterization in Factual Text Involving Named Entities (2022.emnlp-main)

Copied to clipboard

Challenge: Existing topic-based novelty detection methods do not perform semantic reasoning involving relations between named entities in text and their background knowledge.
Approach: They propose a model to detect whether a text is novel or not . they propose to use a factual text to characterize novelty.
Outcome: The proposed model outperforms 10 baselines by large margins on the novelty detection task.
Novelty Goes Deep. A Deep Neural Solution To Document Level Novelty Detection (C18-1)

Copied to clipboard

Challenge: Existing methods for document-level novelty detection are limited and do not require manual feature engineering.
Approach: They propose a deep Convolutional Neural Networks based model to classify a document as novel or redundant on the basis of documents already seen by the system.
Outcome: The proposed model outperforms the state-of-the-art on a document-level novelty detection dataset by a margin of 5% in terms of accuracy.
NovAScore: A New Automated Metric for Evaluating Document Level Novelty (2025.coling-main)

Copied to clipboard

Challenge: Recent research has focused on identifying text that introduces new, previously unknown information, but has seen a decline in novelty detection due to the rise of large language models.
Approach: They propose a novel automated metric for evaluating document-level novelty that aggregates the novelty and salience scores of atomic information and provides high interpretability and a detailed analysis of a document's novelty.
Outcome: The proposed metric scores high on the TAP-DLND 1.0 dataset and a human-annotated dataset.
NLP-ADBench: NLP Anomaly Detection Benchmark (2025.findings-emnlp)

Copied to clipboard

Challenge: Anomaly detection (AD) is an important machine learning task, but its effectiveness in detecting harmful content, phishing attempts, and spam reviews is limited.
Approach: They introduce NLP-ADBench, the most comprehensive NLP anomaly detection benchmark to date . it includes eight curated datasets and 19 state-of-the-art algorithms .
Outcome: The NLP-ADBench benchmark includes 19 state-of-the-art methods and 8 curated datasets . no single model dominates across all datasets, indicating need for automated model selection .
DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for document hierarchy parsing are limited due to the small scale and inconsistency of datasets.
Approach: They propose a document hierarchy parsing dataset to compensate for the data scarcity problem and propose 'dHP' framework to grasp fine-grained text content and coarse-grounded pattern at layout element level.
Outcome: The proposed framework grasps both fine-grained text content and coarse-grounded pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling multi-page and multi-level challenges.
A Survey on Natural Language Processing for Fake News Detection (2020.lrec-1)

Copied to clipboard

Challenge: Automated fake news detection is a critical but challenging problem in NLP . social media has accelerated the spread of fake news, threatening public safety .
Approach: They describe the challenges involved in fake news detection and describe related tasks . they outline promising research directions and highlight the difference between fake news and related tasks.
Outcome: The proposed models are more fine-grained, detailed, fair, and practical.
BREAKING! Presenting Fake News Corpus for Automated Fact Checking (P19-2)

Copied to clipboard

Challenge: a new study shows that fake news spreads faster than mainstream articles on the same topic . however, there is no dataset containing compelling fake and questionable news articles .
Approach: They introduce manually verified corpus of compelling fake and questionable news articles on the USA politics . they plan to extend the corpus in the future and use it for automated fake news detection.
Outcome: The proposed model is based on linguistic features and will be extended in the future . it will be used to improve the existing model and improve the tools in the field of fake news detection .
A Unified Evaluation Framework for Novelty Detection and Accommodation in NLP with an Instantiation in Authorship Attribution (2023.findings-acl)

Copied to clipboard

Challenge: State-of-the-art natural language processing models have been shown to achieve remarkable performance in ‘closed-world’ settings where all the labels in the evaluation set are known at training time.
Approach: They propose a multi-stage task to evaluate a system's performance on pipelined novelty ‘detection’ and ‘accommodation’ tasks.
Outcome: The proposed model performs poorly in ‘closed-world’ settings where all the labels in the evaluation set are known at training time.
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions (2024.lrec-main)

Copied to clipboard

Challenge: SciDMT is an enhanced and expanded corpus for scientific mention detection . existing corpora are limited by their small volume and entity linking capabilities .
Approach: They propose to enhance SciDMT, an annotated scientific corpus for scientific mention detection.
Outcome: The proposed corpus is the largest for scientific entity mention detection . it is based on deep learning architectures like SciBERT and GPT-3.5 .
A Corpus of Metaphor Novelty Scores for Syntactically-Related Word Pairs (L18-1)

Copied to clipboard

Challenge: Existing data on metaphor novelty are limited, making it difficult to perform research on this topic.
Approach: They propose to release a corpus of metaphor novelty scores for syntactically related word pairs . they establish a performance benchmark to which future researchers can compare .
Outcome: The proposed corpus of metaphor novelty scores is compared to other datasets . it performs better than chance or nave strategies, the authors show .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations