Challenge: Xu et al., 2020 focus on semi-structured document classification in a zero-shot setting . positional, layout, and style information play a vital role in interpreting such documents .
Approach: They propose a matching-based approach that relies on a pairwise contrastive objective for pretraining and fine-tuning.
Outcome: The proposed method significantly improves Macro F1 in the zero-shot learning setting.

Similar Papers

Label Agnostic Pre-training for Zero-shot Text Classification (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to text classification assume a fixed set of labels . however, in real-world applications, there exists an infinite label space for describing a given text .
Approach: They propose two new methods that inject aspect-level understanding into pre-trained models at train time to improve zero-shot generalization.
Outcome: The proposed methods improve zero-shot generalization on a set of challenging datasets.
Zero-Shot Text Classification with Self-Training (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large pretrained language models have increased attention to zero-shot text classification.
Approach: They propose a plug-and-play method to bridge this gap by requiring only class names along with an unlabeled dataset.
Outcome: The proposed model can be trained on a natural language inference dataset and performs on dozens of unseen tasks without the need for domain expertise or trial and error.
PESCO: Prompt-enhanced Self Contrastive Learning for Zero-shot Text Classification (2023.acl-long)

Copied to clipboard

Challenge: Existing text classification frameworks require large amounts of human-labeled documents to train .
Approach: They propose a contrastive learning framework that improves zero-shot text classification . they add prompts to enhance label retrieval and use retrieved labels to enrich training .
Outcome: The proposed framework achieves state-of-the-art on four benchmark text classification datasets.
Zero-Shot Text Classification via Self-Supervised Tuning (2023.findings-acl)

Copied to clipboard

Challenge: Existing solutions to zero-shot text classification use pre-trained language models or large-scale annotated data.
Approach: They propose a self-supervised learning paradigm to solve zero-shot text classification tasks by tuning the language models with unlabeled data.
Outcome: The proposed model outperforms the state-of-the-art models on 7 out of 10 tasks and is less sensitive to prompt design.
Few-shot Prompting for Pairwise Ranking: An Effective Non-Parametric Retrieval Model (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that training examples improve zero-shot performance of supervised ranking models.
Approach: They propose to augment supervised ranking models with pairs of queries and documents to improve their performance.
Outcome: The proposed model outperforms the unsupervised models on in-domain and out-domain retrieval benchmarks.
A Representation Sharpening Framework for Zero Shot Dense Retrieval (2026.eacl-long)

Copied to clipboard

Challenge: Zero-shot dense retrieval requires generic, pretrained DRs, which struggle to represent semantic differences between similar documents.
Approach: They propose a training-free representation sharpening framework that augments a document’s representation with information that helps differentiate it from similar documents in the corpus.
Outcome: The proposed framework is compatible with prior approaches to zero-shot dense retrieval and consistently improves their performance.
Zero-shot Text Classification via Reinforced Self-training (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to learn from unlabeled data are difficult for zero-shot text classification tasks.
Approach: They propose a self-training based method to efficiently leverage unlabeled data.
Outcome: The proposed method significantly outperforms existing methods in zero-shot text classification tasks on benchmarks and a real-world e-commerce dataset.
ReGen: Zero-Shot Text Classification via Training Data Generation with Progressive Dense Retrieval (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies show that large pretrained language models can generate training data with no task-specific or cross-task data.
Approach: They propose a retrieval-enhanced framework to create training data from a general-domain unlabeled corpus.
Outcome: The proposed framework achieves 4.3% gain over baselines and saves 70% of time compared with baselines using large language models.
Revisiting Document Representations for Large-Scale Zero-Shot Learning (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for visual recognition use visual attributes carefully annotated by humans.
Approach: They propose a semi-automatic mechanism for visual sentence extraction that leverages document section headers and clustering structure of visual sentences.
Outcome: The proposed method improves on the ImageNet dataset with 10,000 unseen classes.
Structural Contrastive Representation Learning for Zero-shot Multi-label Text Classification (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for zero-shot multi-label text classification struggle with accuracy and poor training efficiency.
Approach: They propose a structural contrastive representation learning approach that uses randomized text segmentation to generate high-quality contrastive pairs.
Outcome: The proposed approach improves accuracy and speed up training time on publicly available datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations