Self-Training with Weak Supervision (2021.naacl-main)

Copied to clipboard

Challenge: State-of-the-art deep neural networks require large amounts of labeled training data that is expensive to obtain or not available for many tasks.
Approach: They propose a weak supervision framework that leverages all available data for a given task . they leverage task-specific unlabeled data through self-training with a model that predicts pseudo-labels for instances that may not be covered by weak rules .
Outcome: The proposed framework improves on state-of-the-art datasets on six benchmark tasks.

Similar Papers

Denoising Multi-Source Weak Supervision for Neural Text Classification (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent years have witnessed the rapid development of deep neural networks (DNNs) for text classification problems.
Approach: They propose a label denoiser which estimates the source reliability using a conditional soft attention mechanism and reduces label noise by aggregating rule-annotated weak labels.
Outcome: The proposed model outperforms state-of-the-art methods on sentiment, topic, and relation classifications and achieves comparable performance with fully-supervised methods even without labeled data.
Handling Noisy Labels for Robustly Learning from Self-Training Data for Low-Resource Sequence Labeling (N19-3)

Copied to clipboard

Challenge: In low-resource environments, self-training is less effective due to unreliable annotations . we combine self-teaching with noise handling to clean the self-labeled data .
Approach: They propose to combine self-training with noise handling to clean unlabeled data . they propose to model clean and noisy labels separately to improve performance .
Outcome: The proposed method performs better than baseline methods on Chunking and NER.
Meta Self-Refinement for Robust Learning with Weak Supervision (2023.eacl-main)

Copied to clipboard

Challenge: Recent methods leverage self-training to build noise-resistant models . however, the teacher trained under weak supervision may have fitted a substantial amount of noise and therefore produce incorrect pseudo-labels.
Approach: They propose a framework that encourages teacher to refine its pseudo-labels to effectively combat label noise from weak supervision.
Outcome: The proposed framework outperforms state-of-the-art methods by 11.4% in accuracy and 9.26% in F1 score on eight NLP benchmarks.
Named Entity Recognition through Deep Representation Learning and Weak Supervision (2021.findings-acl)

Copied to clipboard

Challenge: Weakly supervised named entity recognition (NER) uses noisy labels to estimate the true labels of a dataset.
Approach: They propose a model to learn optimal assignments of latent NER tags using observed tokens and weak labels provided by labeling functions.
Outcome: The proposed model improves the quality of weak labels on four public datasets.
META: Metadata-Empowered Weak Supervision for Text Classification (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for weakly supervised text classification use text data alone to generate pseudo-labels . strong label indicators exist in metadata and it has been long overlooked due to challenges .
Approach: They propose a framework that leverages metadata as an additional source of weak supervision by combining text data and metadata into a text-rich network.
Outcome: The proposed framework exploits metadata as an additional source of weak supervision.
Automatic Rule Induction for Efficient Semi-Supervised Learning (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to generalize from labeled and unlabeled data are difficult to explain and behave unreliably.
Approach: They propose a framework for automatic discovery and integration of symbolic rules into pretrained transformer models by using an attention mechanism.
Outcome: The proposed framework can improve state-of-the-art methods with no manual effort and minimal computational overhead.
Leveraging Training Dynamics and Self-Training for Text Classification (2022.findings-emnlp)

Copied to clipboard

Challenge: Semi-supervised learning (SSL) is a promising technique for improving deep learning models when training data is scarce.
Approach: They propose a semi-supervised learning approach that leverages training dynamics of unlabeled data.
Outcome: The proposed method achieves an average increase in F1 score of 3.5% over baselines in low resource settings.
Weaker Than You Think: A Critical Look at Weakly Supervised Learning (2023.acl-long)

Copied to clipboard

Challenge: Weakly supervised learning is a popular approach for training machine learning models in low-resource settings.
Approach: They propose to use weakly supervised learning to train models with noisy labels from weak sources instead of collecting expensive human annotations.
Outcome: The proposed methods outperform weakly supervised methods on various NLP datasets and tasks on the test sets.
Prototype-Representations for Training Data Filtering in Weakly-Supervised Information Extraction (2022.emnlp-industry)

Copied to clipboard

Challenge: Weak supervision and data programming are powerful tools to support information extraction models.
Approach: They propose a prototype-based method to denoise weakly supervised training data . they use a model to model the correct contexts for a given target value .
Outcome: The proposed method achieves 9% accuracy gain in attribute value extraction in e-commerce websites.
Neural Networks Against (and For) Self-Training: Classification with Small Labeled and Large Unlabeled Sets (2023.findings-acl)

Copied to clipboard

Challenge: Existing models for text classification suffer from the semantic drift problem, which is a problem for self-training.
Approach: They propose a semi-supervised text classifier based on self-training using one positive and one negative property of neural networks.
Outcome: The proposed model outperforms ten baseline models in five benchmarks and is additive to language model pretraining.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations