Challenge: Existing approaches to few-shot text classification require domain expertise and an understanding of the language model's abilities to define the mapping between words and labels.
Approach: They propose a method that converts textual inputs to cloze questions that contain some form of task description and processes them with a pretrained language model to map the predicted words to labels.
Outcome: The proposed approach performs almost as well as hand-crafted label-to-word mappings for a number of tasks with small amounts of training data.

Similar Papers

The Benefits of Label-Description Training for Zero-Shot Text Classification (2023.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models have improved zero-shot text classification by allowing the transfer of semantic knowledge from the training data to classify among specific label sets in downstream tasks.
Approach: They propose to use a small finetuning dataset to describe the labels for a task and to use it to further improve zero-shot accuracies.
Outcome: The proposed model is more accurate than zero-shot by 17-19% absolute across topic and sentiment datasets and more robust to choices required for zero- shot classification.
Distinct Label Representations for Few-Shot Text Classification (2021.acl-short)

Copied to clipboard

Challenge: Existing methods for few-shot text classification ignore the semantic relevance of labels and are difficult to train because of the lack of training examples.
Approach: They propose a method that generates distinct label representations that embed information specific to each label.
Outcome: The proposed method significantly improves few-shot text classification across models and datasets.
Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference (2021.eacl-main)

Copied to clipboard

Challenge: Existing approaches to learning from examples are limited due to the vast number of languages, domains and tasks.
Approach: They propose a semi-supervised training procedure that reformulates input examples as cloze-style phrases to help language models understand a given task.
Outcome: The proposed approach outperforms supervised training and strong semi-supervised approaches in low-resource settings by a large margin.
Label Semantics for Few Shot Named Entity Recognition (2022.findings-acl)

Copied to clipboard

Challenge: Named entity recognition (NER) is a fundamental natural language understanding task that requires large amounts of high quality annotated in-domain data.
Approach: They propose a neural architecture that leverages the semantic information in the names of the labels to give the model additional signal and enriched priors.
Outcome: The proposed model is especially effective in low resource settings.
Few-Shot Learning with Siamese Networks and Label Tuning (2022.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that few-shot text classification is a poor solution for training data-intensive tasks.
Approach: They propose a method that embeds texts and labels into classifiers with proper pre-training.
Outcome: The proposed approach reduces inference cost by increasing the number of labels and embeddings.
Constrained Language Models Yield Few-Shot Semantic Parsers (2021.emnlp-main)

Copied to clipboard

Challenge: Large pretrained language models excel at generating natural language, but they are not efficient for task specific semantic parsing.
Approach: They propose to use large pretrained language models as few-shot semantic parsers . they paraphrase inputs into a controlled sublanguage resembling English .
Outcome: The proposed model can generate surprisingly accurate models on multiple tasks with minimal code and data.
Few-Shot and Zero-Shot Multi-Label Learning for Structured Label Spaces (D18-1)

Copied to clipboard

Challenge: Large multi-label datasets contain labels that occur thousands of times (frequent group), those that occur only a few times (few-shot group) and labels that never appear in the training dataset (zero-shot groups).
Approach: They perform a fine-grained evaluation to understand how state-of-the-art methods perform on infrequent labels.
Outcome: The proposed methods improve on two publicly available datasets for multi-label text classification.
Zero-Shot Text Classification with Self-Training (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large pretrained language models have increased attention to zero-shot text classification.
Approach: They propose a plug-and-play method to bridge this gap by requiring only class names along with an unlabeled dataset.
Outcome: The proposed model can be trained on a natural language inference dataset and performs on dozens of unseen tasks without the need for domain expertise or trial and error.
Human in the Loop: How to Effectively Create Coherent Topics by Manually Labeling Only a Few Documents per Class (2024.lrec-main)

Copied to clipboard

Challenge: Few-shot methods for accurate modeling under sparse label-settings are still challenging in document classification.
Approach: They propose to combine supervised few-shot learning with a topic extraction method to generate coherent topics in large text corpora.
Outcome: The proposed method outperforms unsupervised topic modeling methods in document classification.
Label-Aware Automatic Verbalizer for Few-Shot Text Classification in Mid-To-Low Resource Languages (2024.acl-srw)

Copied to clipboard

Challenge: Prompt-based learning has shown its effectiveness in few-shot text classification.
Approach: They propose a prompt-based learning verbalizer that automatically selects a word to represent each class . they use the label name along with the conjunction "and" to induce the model to generate more effective words for the verbaliser.
Outcome: The proposed method outperforms existing verbalizers on four Southeast Asian languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations