Label Noise in Context (2020.acl-demos)

Copied to clipboard

Challenge: Label noise—incorrectly or ambiguously labeled training examples—can negatively impact model performance.
Approach: They propose a noise-detection method that uses an example's neighborhood within the training set to reduce false positives and provide an explanation as to why the ex ample was flagged as noise.
Outcome: The proposed method outperforms the state-of-the-art on precision and F0.5-score on short-text classification datasets.

Similar Papers

NoisywikiHow: A Benchmark for Learning with Real-world Noisy Labels in Natural Language Processing (2023.findings-acl)

Copied to clipboard

Challenge: Large-scale datasets in the real world often contain label noise, which can cause model overfitting and degrade generalization.
Approach: They propose to use label noise to imitate human errors in annotations . they use a noisy label noise benchmark to evaluate their methods .
Outcome: The proposed benchmarks are different from data with heterogeneous label noises in the real world.
Noisy Label Regularisation for Textual Regression (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to regularise noisy labels are ineffective in the face of noisy data.
Approach: They propose a method that regularises noisy labels and prevents error propagation from the input layer.
Outcome: The proposed method regularises noisy labels and improves generalisation performance over real-world human-disagreement annotations and randomly-corrupted and data-augmented labels.
Adaptive Textual Label Noise Learning based on Pre-trained Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to learning with noisy labels are limited due to the time and labor costs involved.
Approach: They propose an adaptive warm-up and hybrid training frameworks to learn with noisy labels based on pre-trained models.
Outcome: The proposed approach performs comparable or even surpasses state-of-the-art methods in various noise scenarios, including scenarios with the mixture of multiple types of noise.
Hide and Seek in Noise Labels: Noise-Robust Collaborative Active Learning with LLMs-Powered Assistance (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for learning from noisy labels are difficult to improve . existing methods identify noisy labels and use active learning to query experts .
Approach: They propose a collaborative learning framework to combine LLMs and small models for learning from noisy labels.
Outcome: The proposed framework outperforms state-of-the-art baselines on synthetic and real-world noise datasets.
Feature-Dependent Confusion Matrices for Low-Resource NER Labeling with Noisy Labels (D19-1)

Copied to clipboard

Challenge: Existing approaches to improve supervised labeling with noisy training data do not take the input features into account or they need to learn the noise modeling from scratch.
Approach: They propose to cluster training data using input features and compute different confusion matrices for each cluster.
Outcome: The proposed model improves on low-resource named entity recognition settings in several languages, compared with other models which do not take the input features into account or need to learn noise modeling from scratch.
Enhancing Multi-Label Text Classification under Label-Dependent Noise: A Label-Specific Denoising Framework (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing noisy multi-label text classification methods rely on the class-conditional noise assumption, but in practice, noisy labels exhibit a certain degree of correlation with the true labels.
Approach: They propose a label-specific denoising framework to counteract label-dependent noise by evaluating loss information, ranking information, and feature centroid.
Outcome: The proposed framework significantly improves over existing state-of-the-art models under both synthetic and real-world noise conditions.
Learning on Imbalanced Noisy Data via Debiased Sample Selection and LLM-Driven Annotation (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to learning with noisy labels are prone to selection bias and training bias . obtaining large-scale high-quality datasets is expensive and time-consuming in practical scenarios .
Approach: They propose an imbalanced learning with noisy labels task to let model learn from noisy labels . they first conduct debiased sample selection to better separate clean samples from noisy samples . then they feed selected clean samples to active annotator large language models for re-annotating noisy samples.
Outcome: The proposed method is superior to existing methods on synthetic and real-world datasets.
An Effective Label Noise Model for DNN Text Classification (N19-1)

Copied to clipboard

Challenge: Existing methods to train deep neural networks with label noise are limited to image classification models . label noise is important because of the large number of errors and errors in training datasets .
Approach: They propose a non-linear processing layer that models label noise into a convolutional neural network (CNN) they add a noise model layer on top of their target model to account for label noise .
Outcome: The proposed approach is robust to label noise and can learn better sentences . it is based on extensive experiments on text classification datasets .
Noise-Robust Fine-Tuning of Pretrained Language Models via External Guidance (2023.findings-emnlp)

Copied to clipboard

Challenge: Pretrained Language Models (PLMs) are advanced but data labels are noisy due to the complex annotation process.
Approach: They propose a framework for fine-tuning PLMs using noisy labels that incorporates guidance from Large Language Models like ChatGPT.
Outcome: Experiments on synthetic and real-world noisy datasets show that the proposed framework outperforms the state-of-the-art framework.
Detecting Label Errors by Using Pre-Trained Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for label error detection focus on label errors in training data.
Approach: They propose a method for introducing realistic, human-originated label noise into existing crowdsourced datasets such as SNLI and TweetNLP.
Outcome: The proposed method outperforms existing methods for detecting label errors in natural language datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations