Challenge: Existing approaches to named entity recognition (NER) in Chinese are limited by the lack of annotated data.
Approach: They propose a method which can automatically populate annotated training data without humancost by using distant supervision.
Outcome: The proposed method performs better than comparison systems on two datasets.

Similar Papers

Reinforcement-based denoising of distantly supervised NER with partial annotation (D19-61)

Copied to clipboard

Challenge: Existing named entity recognition systems rely on large amounts of human-labeled data for supervision, but the result is noisy.
Approach: They propose to use partial annotation to address false negative cases and implement a reinforcement learning strategy to identify false positive instances.
Outcome: The proposed model reduces the amount of manually annotated data required to perform NER in a new domain.
Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-Training (2021.emnlp-main)

Copied to clipboard

Challenge: Named entity recognition models require abundant high-quality annotations to train . distant supervision may induce incomplete and noisy labels, making supervised learning ineffective.
Approach: They propose a noise-robust learning scheme for training named entity recognition models using only distantly-labeled data and a self-training method that uses contextualized augmentations created by pre-trained language models.
Outcome: The proposed method outperforms existing supervised NER models on three datasets by significant margins.
Learning Named Entity Tagger using Domain-Specific Dictionary (D18-1)

Copied to clipboard

Challenge: Existing methods to build reliable named entity recognition systems require large amounts of manually-labeled training data.
Approach: They propose a revised fuzzy CRF layer to handle tokens with multiple possible labels to address noisy distant supervision.
Outcome: The proposed model can handle tokens with multiple possible labels under the traditional framework and improves on the existing model with a new Tie or Break scheme.
Named Entity Recognition without Labelled Data: A Weak Supervision Approach (2020.acl-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) performance often degrades when applied to target domains that differ from the texts observed during training.
Approach: They propose a method to learn NER models in the absence of labelled data through weak supervision by using a broad spectrum of labelling functions to automatically annotate texts from the target domain.
Outcome: The proposed approach improves on two English datasets and shows that it improves by 7 percentage points on entity-level F1 scores compared to an out-of-domain neural NER model.
Better Modeling of Incomplete Annotations for Named Entity Recognition (N19-1)

Copied to clipboard

Challenge: Existing approaches to named entity recognition (NER) assume that the training data is fully annotated with named entity information.
Approach: They propose a supervised setup for named entity recognition where annotated data is assumed to be available during training.
Outcome: The proposed approach is able to recognize named entities with incomplete annotations.
Distantly Supervised Named Entity Recognition via Confidence-Based Multi-Class Positive and Unlabeled Learning (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for named entity recognition suffer from incomplete annotations due to incompleteness of external knowledge bases.
Approach: They propose a method to solve the named entity recognition problem under distant supervision using dictionaries and knowledge bases.
Outcome: The proposed method outperforms existing methods on two benchmark datasets labeled by various knowledge bases.
Label Refinement via Contrastive Learning for Distantly-Supervised Named Entity Recognition (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods to locate and classify entities using knowledge bases and unlabeled corpus are expensive and limited application.
Approach: They propose to use a method to directly learn the distant label refinement knowledge by imitating annotations of different qualities and comparing them in contrastive learning frameworks.
Outcome: The proposed method can give modified suggestions on distant data without additional supervised labels and thus reduces the requirement on the quality of the knowledge bases.
Distantly Supervised Named Entity Recognition using Positive-Unlabeled Learning (P19-1)

Copied to clipboard

Challenge: Empirical studies on four public NER datasets demonstrate the effectiveness of our proposed method.
Approach: They propose a method to perform named entity recognition using unlabeled data and named entity dictionaries.
Outcome: The proposed method can estimate task loss as if there is fully labeled data.
Self-Cleaning: Improving a Named Entity Recognizer Trained on Noisy Data with a Few Clean Instances (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods to train named entity recognition models on noisy data are expensive and time-intensive to accumulate.
Approach: They propose to denoise noisy NER data with guidance from a small set of clean instances.
Outcome: The proposed method can improve on large-scale datasets with a small guidance set.
Noise-Robust Training with Dynamic Loss and Contrastive Learning for Distantly-Supervised Named Entity Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Named entity recognition (NER) is a task in natural language processing that aims at locating entity mentions in a given sentence and assigning them to certain types.
Approach: They propose to use a dynamic loss function to better adapt to the changing noise during the training process and incorporate token level contrastive learning to fully utilize the noisy data.
Outcome: The proposed method outperforms existing NER models on three benchmark datasets and outperformed existing models by significant margins.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations