Challenge: Existing methods for document-level novelty detection are limited and do not require manual feature engineering.
Approach: They propose a deep Convolutional Neural Networks based model to classify a document as novel or redundant on the basis of documents already seen by the system.
Outcome: The proposed model outperforms the state-of-the-art on a document-level novelty detection dataset by a margin of 5% in terms of accuracy.

Similar Papers

TAP-DLND 1.0 : A Corpus for Document Level Novelty Detection (L18-1)

Copied to clipboard

Challenge: Detecting novelty of an entire document is an AI frontier problem . present state-of-the-art text matching techniques are unable to process such redundancy.
Approach: They propose a document-level novelty detection resource that can be used to benchmark techniques . they crawl news documents across several domains and use it to find out whether they contain new information .
Outcome: The proposed dataset is compared with a standard system for document novelty detection . the proposed system can detect elements that have not appeared before, or new or original .
Semantic Novelty Detection and Characterization in Factual Text Involving Named Entities (2022.emnlp-main)

Copied to clipboard

Challenge: Existing topic-based novelty detection methods do not perform semantic reasoning involving relations between named entities in text and their background knowledge.
Approach: They propose a model to detect whether a text is novel or not . they propose to use a factual text to characterize novelty.
Outcome: The proposed model outperforms 10 baselines by large margins on the novelty detection task.
Rethinking Complex Neural Network Architectures for Document Classification (N19-1)

Copied to clipboard

Challenge: Neural network models for many NLP tasks have grown increasingly complex in recent years . authors of recent papers question the necessity of such architectures and find them quite effective .
Approach: They propose to use regularization techniques borrowed from language modeling to improve model accuracy . they find that a simple biLSTM architecture with appropriate regularization yields competitive results .
Outcome: a simple biLSTM model outperforms the state-of-the-art on four benchmark datasets . authors say that improvements are not real, but are attributed to mundane reasons .
Deep Adversarial Learning for NLP (N19-5)

Copied to clipboard

Challenge: Adversarial learning is a game-theoretic learning paradigm that has achieved huge successes in the field of Computer Vision recently.
Approach: This tutorial introduces the foundations of deep adversarial learning and some practical problems and solutions in NLP.
Outcome: This tutorial introduces the foundations of deep adversarial learning and some practical problems and solutions in NLP.
Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019) (D19-61)

Copied to clipboard

Challenge: EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China .
Approach: EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China . call for papers for this second workshop met with a strong response .
Outcome: the EMNLP-IJCNLP 2019 workshop on deep learning approaches for low-resource natural language processing takes place in Hong Kong, China.
Convolutional Neural Networks with Recurrent Neural Filters (D18-1)

Copied to clipboard

Challenge: Convolutional neural networks (CNNs) use recurrent neural networks as convolution filters to capture language compositionality and long-term dependencies.
Approach: They propose to use recurrent neural networks (RNNs) as convolution filters to capture language compositionality and long-term dependencies.
Outcome: The proposed convolutional neural networks achieve state-of-the-art on two sentences and the Stanford Sentiment Treebank.
Threat Scenarios and Best Practices to Detect Neural Fake News (2022.coling-1)

Copied to clipboard

Challenge: During the COVID-19 pandemic, inaccurate information made it hard for people to find reliable guidance when they needed it.
Approach: They propose to use pretrained language models to generate fluent, original text . they argue that strong detectors should be released along with new generators .
Outcome: The proposed system is prone to shortcut learning and should be released along with new generators.
Representation Learning for Information Extraction from Form-like Documents (2020.acl-main)

Copied to clipboard

Challenge: Form-like documents like invoices, purchase orders, tax forms and insurance quotes are common in day-to-day business workflows, but current techniques for processing them largely still employ manual effort or brittle and error-prone heuristics for extraction.
Approach: They propose an extraction system that uses knowledge of the types of the target fields to generate extraction candidates and a neural network architecture that learns a dense representation of each candidate based on neighboring words in the document.
Outcome: The proposed system generates extraction candidates based on neighboring words in the document and is interpretable, as shown using loss cases.
Neural Models for Documents with Metadata (P18-1)

Copied to clipboard

Challenge: specialized models are often used to model text corpora without metadata . specialized algorithms are not widely used in the digital humanities and political science fields .
Approach: They propose a general neural framework based on topic models to enable customization of metadata.
Outcome: The proposed framework achieves strong performance with a manageable tradeoff between perplexity, coherence, and sparsity.
Semantic Novelty Detection in Natural Language Descriptions (2021.emnlp-main)

Copied to clipboard

Challenge: Existing novelty detection algorithms are coarse-grained, working at the document or topic level.
Approach: They propose to use a fine-grained semantic novelty detection problem to solve a novel novel scene problem.
Outcome: The proposed model outperforms baseline models on the proposed task by large margins.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations