Challenge: Existing methods for weakly-supervised text classification use only class names as supervision . Existing approaches to classify texts without labeled data have significant flaws, including zero-shot instability and context-dependent ambiguities.
Approach: They propose to use wordsets to generate pseudo-labels for unlabeled texts . they propose to train the classifier using a hybrid learning strategy called sync-denoising .
Outcome: The proposed method outperforms all existing prompt and seed methods on 11 datasets by an impressive average of 8 points.

Similar Papers

A Benchmark on Extremely Weakly Supervised Text Classification: Reconcile Seed Matching and Prompting Approaches (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for XWS-TC rely on minimal human guidance . X-WS-tc methods require no humanannotated datasets .
Approach: They propose a benchmarking method to compare two approaches to XWS-TC . they use seed-matching and prompting a language model with instructions to decode label words .
Outcome: The proposed methods are more tolerant to human guidance and more robust to model-based methods.
X-Class: Text Classification with Extremely Weak Supervision (2021.naacl-main)

Copied to clipboard

Challenge: Weak supervision is a problem in text classification, but it requires corpusspecific knowledge.
Approach: They propose a framework for extremely weak supervision that can be used to train a text classifier.
Outcome: The proposed framework outperforms seed-driven weakly supervised methods on 7 benchmark datasets.
Contextualized Weak Supervision for Text Classification (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for weakly supervised text classification generate pseudo-labels in a context-free manner, thus, the ambiguous, context-dependent nature of human language has been long overlooked.
Approach: They propose a framework that provides contextualized weak supervision for text classification . they leverage contextualized representations of word occurrences and seed word information .
Outcome: The proposed framework provides contextualized weak supervision for text classification . it leverages representations of word occurrences and seed word information to differentiate interpretations . the proposed framework also disambiguates initial seed words, making it fully contextualized .
PIEClass: Weakly-Supervised Text Classification with Prompting and Noise-Robust Iterative Ensemble Training (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for text classification use label names of target classes as the only supervision.
Approach: They propose a method that uses keyword-based keyword matching to generate pseudo labels . they propose 'pieclass' module that iteratively trains classifiers and updates pseudo labels.
Outcome: The proposed method achieves better performance than existing strong baselines on seven benchmark datasets and similar performance to fully-supervised classifiers on sentiment classification tasks.
LIME: Weakly-Supervised Text Classification without Seeds (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches to weakly-supervised text classification use only label names as sources of supervision.
Approach: They propose a framework for weakly-supervised text classification that replaces seed-word generation with entailment-based pseudo-classification.
Outcome: The proposed framework outperforms baselines and state-of-the-art in 4 benchmarks.
Denoising Multi-Source Weak Supervision for Neural Text Classification (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent years have witnessed the rapid development of deep neural networks (DNNs) for text classification problems.
Approach: They propose a label denoiser which estimates the source reliability using a conditional soft attention mechanism and reduces label noise by aggregating rule-annotated weak labels.
Outcome: The proposed model outperforms state-of-the-art methods on sentiment, topic, and relation classifications and achieves comparable performance with fully-supervised methods even without labeled data.
Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text Classification (2024.emnlp-main)

Copied to clipboard

Challenge: Recent work generates pseudo labels by mining texts similar to the class names from the raw corpus, but there is a high risk that LLMs cannot generate in-distribution data, leading to ungeneralizable classifiers.
Approach: They propose to use LLMs to generate pseudo labels by mining masked templates from corpus . they then use state-of-the-art LLM to synthesize near-distribution texts falling into minority classes .
Outcome: The proposed framework improves on the previous methods for extremely weak-supervised text classification.
META: Metadata-Empowered Weak Supervision for Text Classification (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for weakly supervised text classification use text data alone to generate pseudo-labels . strong label indicators exist in metadata and it has been long overlooked due to challenges .
Approach: They propose a framework that leverages metadata as an additional source of weak supervision by combining text data and metadata into a text-rich network.
Outcome: The proposed framework exploits metadata as an additional source of weak supervision.
Debiasing Made State-of-the-art: Revisiting the Simple Seed-based Weak Supervision for Text Classification (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in weakly supervised text classification focus on designing sophisticated methods to turn high-level human heuristics into quality pseudo-labels.
Approach: They propose to use a seed matching-based method to generate quality pseudo-labels by deleting the seed words present in the matched input text.
Outcome: The proposed method can be improved significantly by deleting the seed words in the matched input text with a high deletion ratio.
Seed Word Selection for Weakly-Supervised Text Classification with Unsupervised Error Estimation (2021.naacl-srw)

Copied to clipboard

Challenge: Weakly-supervised text classification aims to induce text classifiers from only a handful of user-provided seed words.
Approach: They propose to use user-provided seed words to induce text classifiers using only a handful of carefully chosen seed words.
Outcome: The proposed method outperforms baseline model using only category name seed words and achieves comparable performance as a counterpart using expert-annotated seed words.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations