Challenge: Existing studies develop effective pseudo-labeling methods, but they struggle with unlabeled data that have imbalanced classes mismatched with the labeled data.
Approach: They propose to use pseudo-labeling to train text classification models with few labeled data and massive unlabeled data.
Outcome: Empirical results show that the proposed model outperforms state-of-the-art methods on 3 common benchmarks.

Similar Papers

Prototype-Guided Pseudo Labeling for Semi-Supervised Text Classification (2023.acl-long)

Copied to clipboard

Challenge: Existing semi-supervised text classification methods suffer from categorical boundary issues . existing methods suffer by ambiguous categoric boundaries, making it difficult to generate reliable pseudo-labels for each category.
Approach: They propose a semi-supervised framework that assigns pseudo-labels to unlabeled data . they exploit categorical prototypes to assimilate instance representations within the same category .
Outcome: Empirical studies show that the proposed framework is effective . it uses prototypical cluster separation and prototypical-center data selection .
Semi-Supervised Text Classification with Balanced Deep Representation Distributions (2021.acl-long)

Copied to clipboard

Challenge: Semi-Supervised Text Classification (SSTC) is a type of self-training that uses labeled and unlabeled data to perform certain applications.
Approach: They propose a method to initialize a deep classifier by training over labeled texts . they then alternatively predict unlabeled texts as their pseudo-labels and train them over the mixture .
Outcome: Empirical results show that the proposed method is more accurate when labeled texts are scarce.
Rank-Aware Negative Training for Semi-Supervised Text Classification (2023.tacl-1)

Copied to clipboard

Challenge: Semi-supervised text classification-based paradigms employ the spirit of self-training, but the accuracy of pseudo-labels can be a problem in real-world scenarios.
Approach: They propose a Rank-aware Negative Training framework to address SSTC in noisy label learning . they rank unlabeled texts based on evidential support from the labeled texts.
Outcome: The proposed framework overcomes state-of-the-art alternatives and achieves competitive performance in other scenarios.
JointMatch: A Unified Approach for Diverse and Collaborative Pseudo-Labeling to Semi-Supervised Text Classification (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to semi-supervised text classification suffer from pseudo-label bias and error accumulation.
Approach: They propose a pseudo-labeling approach to semi-supervised text classification that unifies ideas from semi-semi-supervised learning and the task of learning with noise.
Outcome: The proposed approach achieves a significant improvement on benchmark datasets even in the extremely-scarce-label setting.
Neural Networks Against (and For) Self-Training: Classification with Small Labeled and Large Unlabeled Sets (2023.findings-acl)

Copied to clipboard

Challenge: Existing models for text classification suffer from the semantic drift problem, which is a problem for self-training.
Approach: They propose a semi-supervised text classifier based on self-training using one positive and one negative property of neural networks.
Outcome: The proposed model outperforms ten baseline models in five benchmarks and is additive to language model pretraining.
Joint Speech Transcription and Translation: Pseudo-Labeling with Out-of-Distribution Data (2023.findings-acl)

Copied to clipboard

Challenge: a recent study shows that self-training can improve upon fully supervised baselines in low-resource settings for several sequence-to-sequence tasks.
Approach: They propose to use pseudo-labeling to label unsupervised data and add it to the training pool.
Outcome: The proposed setup improves on the unsupervised data by using pseudo-labeling . the proposed setup provides 0.4% absolute WER and 2.1 BLEU points for En–De .
Robust Representation Learning with Reliable Pseudo-labels Generation via Self-Adaptive Optimal Transport for Short Text Clustering (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to short text clustering are prone to degenerate solutions and noisy data.
Approach: They propose a model to improve robustness against imbalanced and noisy data . they propose self-adaptive optimal transport and class-wise contrastive learning .
Outcome: The proposed model outperforms the state-of-the-art models on eight short text clustering datasets.
Unraveling the Dynamics of Semi-Supervised Hate Speech Detection: The Impact of Unlabeled Data Characteristics and Pseudo-Labeling Strategies (2024.findings-eacl)

Copied to clipboard

Challenge: Semi-supervised learning addresses the need for large amounts of labeled training data for state-of-the-art approaches.
Approach: They propose to leverage unlabeled data to reduce the amount of annotated data required for machine learning based hate speech detection by using a semi-supervised approach.
Outcome: The proposed approach reduces the amount of annotated data required by state-of-the-art models by leveraging unlabeled data.
STSPL-SSC: Semi-Supervised Few-Shot Short Text Clustering with Semantic text similarity Optimized Pseudo-Labels (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for obtaining task-specific labels require prior knowledge of clustering categories and uncontrollable clustering centers.
Approach: They propose a framework for supervised clustering using a discrete process and a robust Contrastive Learning module.
Outcome: The proposed framework outperforms state-of-the-art models on a real-world dataset with just one label per class . the proposed framework is based on k-means clustering and a robust Contrastive Learning module .
Leveraging Training Dynamics and Self-Training for Text Classification (2022.findings-emnlp)

Copied to clipboard

Challenge: Semi-supervised learning (SSL) is a promising technique for improving deep learning models when training data is scarce.
Approach: They propose a semi-supervised learning approach that leverages training dynamics of unlabeled data.
Outcome: The proposed method achieves an average increase in F1 score of 3.5% over baselines in low resource settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations