Extremely Weakly-supervised Text Classification with Wordsets Mining and Sync-Denoising (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for weakly-supervised text classification use only class names as supervision . Existing approaches to classify texts without labeled data have significant flaws, including zero-shot instability and context-dependent ambiguities. |
| Approach: | They propose to use wordsets to generate pseudo-labels for unlabeled texts . they propose to train the classifier using a hybrid learning strategy called sync-denoising . |
| Outcome: | The proposed method outperforms all existing prompt and seed methods on 11 datasets by an impressive average of 8 points. |
Similar Papers
A Benchmark on Extremely Weakly Supervised Text Classification: Reconcile Seed Matching and Prompting Approaches (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for XWS-TC rely on minimal human guidance . X-WS-tc methods require no humanannotated datasets . |
| Approach: | They propose a benchmarking method to compare two approaches to XWS-TC . they use seed-matching and prompting a language model with instructions to decode label words . |
| Outcome: | The proposed methods are more tolerant to human guidance and more robust to model-based methods. |
X-Class: Text Classification with Extremely Weak Supervision (2021.naacl-main)
Copied to clipboard
| Challenge: | Weak supervision is a problem in text classification, but it requires corpusspecific knowledge. |
| Approach: | They propose a framework for extremely weak supervision that can be used to train a text classifier. |
| Outcome: | The proposed framework outperforms seed-driven weakly supervised methods on 7 benchmark datasets. |
Contextualized Weak Supervision for Text Classification (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for weakly supervised text classification generate pseudo-labels in a context-free manner, thus, the ambiguous, context-dependent nature of human language has been long overlooked. |
| Approach: | They propose a framework that provides contextualized weak supervision for text classification . they leverage contextualized representations of word occurrences and seed word information . |
| Outcome: | The proposed framework provides contextualized weak supervision for text classification . it leverages representations of word occurrences and seed word information to differentiate interpretations . the proposed framework also disambiguates initial seed words, making it fully contextualized . |
PIEClass: Weakly-Supervised Text Classification with Prompting and Noise-Robust Iterative Ensemble Training (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for text classification use label names of target classes as the only supervision. |
| Approach: | They propose a method that uses keyword-based keyword matching to generate pseudo labels . they propose 'pieclass' module that iteratively trains classifiers and updates pseudo labels. |
| Outcome: | The proposed method achieves better performance than existing strong baselines on seven benchmark datasets and similar performance to fully-supervised classifiers on sentiment classification tasks. |
LIME: Weakly-Supervised Text Classification without Seeds (2022.coling-1)
Copied to clipboard
| Challenge: | Existing approaches to weakly-supervised text classification use only label names as sources of supervision. |
| Approach: | They propose a framework for weakly-supervised text classification that replaces seed-word generation with entailment-based pseudo-classification. |
| Outcome: | The proposed framework outperforms baselines and state-of-the-art in 4 benchmarks. |
Denoising Multi-Source Weak Supervision for Neural Text Classification (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent years have witnessed the rapid development of deep neural networks (DNNs) for text classification problems. |
| Approach: | They propose a label denoiser which estimates the source reliability using a conditional soft attention mechanism and reduces label noise by aggregating rule-annotated weak labels. |
| Outcome: | The proposed model outperforms state-of-the-art methods on sentiment, topic, and relation classifications and achieves comparable performance with fully-supervised methods even without labeled data. |
Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text Classification (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent work generates pseudo labels by mining texts similar to the class names from the raw corpus, but there is a high risk that LLMs cannot generate in-distribution data, leading to ungeneralizable classifiers. |
| Approach: | They propose to use LLMs to generate pseudo labels by mining masked templates from corpus . they then use state-of-the-art LLM to synthesize near-distribution texts falling into minority classes . |
| Outcome: | The proposed framework improves on the previous methods for extremely weak-supervised text classification. |
META: Metadata-Empowered Weak Supervision for Text Classification (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for weakly supervised text classification use text data alone to generate pseudo-labels . strong label indicators exist in metadata and it has been long overlooked due to challenges . |
| Approach: | They propose a framework that leverages metadata as an additional source of weak supervision by combining text data and metadata into a text-rich network. |
| Outcome: | The proposed framework exploits metadata as an additional source of weak supervision. |
Debiasing Made State-of-the-art: Revisiting the Simple Seed-based Weak Supervision for Text Classification (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in weakly supervised text classification focus on designing sophisticated methods to turn high-level human heuristics into quality pseudo-labels. |
| Approach: | They propose to use a seed matching-based method to generate quality pseudo-labels by deleting the seed words present in the matched input text. |
| Outcome: | The proposed method can be improved significantly by deleting the seed words in the matched input text with a high deletion ratio. |
Seed Word Selection for Weakly-Supervised Text Classification with Unsupervised Error Estimation (2021.naacl-srw)
Copied to clipboard
| Challenge: | Weakly-supervised text classification aims to induce text classifiers from only a handful of user-provided seed words. |
| Approach: | They propose to use user-provided seed words to induce text classifiers using only a handful of carefully chosen seed words. |
| Outcome: | The proposed method outperforms baseline model using only category name seed words and achieves comparable performance as a counterpart using expert-annotated seed words. |