Challenge: Recent work generates pseudo labels by mining texts similar to the class names from the raw corpus, but there is a high risk that LLMs cannot generate in-distribution data, leading to ungeneralizable classifiers.
Approach: They propose to use LLMs to generate pseudo labels by mining masked templates from corpus . they then use state-of-the-art LLM to synthesize near-distribution texts falling into minority classes .
Outcome: The proposed framework improves on the previous methods for extremely weak-supervised text classification.

Similar Papers

X-Class: Text Classification with Extremely Weak Supervision (2021.naacl-main)

Copied to clipboard

Challenge: Weak supervision is a problem in text classification, but it requires corpusspecific knowledge.
Approach: They propose a framework for extremely weak supervision that can be used to train a text classifier.
Outcome: The proposed framework outperforms seed-driven weakly supervised methods on 7 benchmark datasets.
META: Metadata-Empowered Weak Supervision for Text Classification (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for weakly supervised text classification use text data alone to generate pseudo-labels . strong label indicators exist in metadata and it has been long overlooked due to challenges .
Approach: They propose a framework that leverages metadata as an additional source of weak supervision by combining text data and metadata into a text-rich network.
Outcome: The proposed framework exploits metadata as an additional source of weak supervision.
Improving Text Embeddings with Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for obtaining text embeddings require complex training pipelines . authors leverage proprietary LLMs to generate diverse synthetic data for text embeds based on 93 languages .
Approach: They propose a method for obtaining high-quality text embeddings using only synthetic data and less than 1k training steps.
Outcome: The proposed method achieves strong performance on competitive text embedding benchmarks without using any labeled data.
Extremely Weakly-supervised Text Classification with Wordsets Mining and Sync-Denoising (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for weakly-supervised text classification use only class names as supervision . Existing approaches to classify texts without labeled data have significant flaws, including zero-shot instability and context-dependent ambiguities.
Approach: They propose to use wordsets to generate pseudo-labels for unlabeled texts . they propose to train the classifier using a hybrid learning strategy called sync-denoising .
Outcome: The proposed method outperforms all existing prompt and seed methods on 11 datasets by an impressive average of 8 points.
Weakly Supervised Text Classification using Supervision Signals from a Language Model (2022.findings-naacl)

Copied to clipboard

Challenge: Existing weakly supervised text classification methods require a large number of annotated data and human annotations are expensive.
Approach: They propose to query a masked language model with cloze style prompts to obtain supervision signals.
Outcome: The proposed method outperforms baseline methods on three datasets by 2%, 4%, and 3%.
Text Augmentation Using Dataset Reconstruction for Low-Resource Classification (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for text classification use labeled data, but labeles are expensive and difficult to obtain.
Approach: They propose a novel method of data augmentation using the text-generation capabilities of language models.
Outcome: The proposed method improves the current state-of-the-art methods for data augmentation on multi-class datasets.
Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self-Training Approach (2021.naacl-main)

Copied to clipboard

Challenge: Fine-tuned pre-trained language models (LMs) have enormous success in many natural language processing tasks, but they still require excessive labeled data in the fine-tuning stage.
Approach: They propose a framework to enable fine-tuning pre-trained language models with weak supervision without any labeled data.
Outcome: The proposed framework outperforms the strongest baseline and achieves competitive performance with fully-supervised fine-tuning methods.
Out-of-Distribution Detection via LLM-Guided Outlier Generation for Text-attributed Graph (2025.findings-acl)

Copied to clipboard

Challenge: Text-Attributed Graphs (TAGs) are widely used in the real world.
Approach: They propose to use Large Language Models to generate OOD-nodes with high quality . they also use LLMs to integrate existing nodes with LLM-generated edges .
Outcome: The proposed method performs well on samples outside the In-Distribution (ID) data, but it is difficult to obtain high-quality OOD samples in the real world.
Unsupervised Cross-Lingual Representation Learning (P19-4)

Copied to clipboard

Challenge: a comprehensive survey of cutting-edge weakly-supervised and unsupervised cross-lingual word representations is presented .
Approach: This tutorial provides a comprehensive survey of recent work on weakly-supervised and unsupervised cross-lingual word representations.
Outcome: This tutorial provides a comprehensive survey of cutting-edge weakly-supervised and unsupervised word representations.
Investigating Ensemble Methods for Model Robustness Improvement of Text Classifiers (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to reduce model's reliance on bias features ignore the learnability of these features.
Approach: They propose to reduce models' reliance on bias features by first training models with fixed low-capacity models which ignore the learnability of the bias features.
Outcome: The proposed models can perform better on out-of-distribution datasets than baseline models with a more sophisticated model design.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations