Challenge: Pre-trained language models provide the foundations for state-of-the-art performance across a wide range of natural language processing tasks, including text classification.
Approach: They compare the performance of a linear classifier based on word embeddings with a pre-trained language model, i.e., BERT, across a wide range of datasets and classification tasks.
Outcome: The proposed method outperforms baselines in standard datasets with large training sets, but in settings with small training datasets it performs better.

Similar Papers

Give your Text Representation Models some Love: the Case for Basque (2020.lrec-1)

Copied to clipboard

Challenge: Word embeddings and pre-trained language models are expensive to train and are often used by small companies and research groups to build their own.
Approach: They propose to use word embeddings and pre-trained language models to build rich representations of text and improve NLP tasks.
Outcome: The proposed models perform better than publicly available versions in downstream NLP tasks for Basque.
Linear Classifier: An Often-Forgotten Baseline for Text Classification (2023.acl-short)

Copied to clipboard

Challenge: Large-scale pre-trained language models such as BERT are popular solutions for text classification.
Approach: They argue that large-scale pre-trained language models such as BERT are popular solutions for text classification . authors argue that running a simple baseline like linear classifiers on bag-of-words features is important for text classification .
Outcome: The proposed approach may only sometimes get satisfactory results for some problems.
Fusing Label Embedding into BERT: An Efficient Improvement for Text Classification (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to improve text classification performance of pre-trained models have been used to improve their performance.
Approach: They propose a method for improving BERT's performance by using a label embedding technique while keeping almost the same computational cost.
Outcome: The proposed method improves BERT's performance on six text classification benchmark datasets while keeping almost the same computational cost.
On the importance of pre-training data volume for compact language models (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in language modeling have led to computationally intensive and resource-demanding state-of-the-art models.
Approach: They investigate the impact of pre-training data volume on compact language models . they use a French question answering task to train models with as little as 100 MB of text .
Outcome: The results show that pre-training data volume can improve models with as little as 100 MB of text . the results suggest that the model performance is poorer with less data than with larger datasets .
Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks (2020.acl-main)

Copied to clipboard

Challenge: Language models prerained on text from a wide variety of sources form the foundation of today’s NLP.
Approach: They propose to tailor a pretrained model to the domain of a target task by using domain-adaptive pretraining in-domain.
Outcome: The proposed model can be tailored to the domain of a target task and perform well under both high- and low-resource settings.
Text Classification with Few Examples using Controlled Generalization (N19-1)

Copied to clipboard

Challenge: Current training data for text classification is limited, resulting in limited generalization capacity.
Approach: They propose a feed-forward network that can generalize from unlabeled parsed corpora to produce task-specific semantic vectors.
Outcome: The proposed approach is especially effective in low-data scenarios compared to state-of-the-art methods.
Cost-effective Selection of Pretraining Data: A Case Study of Pretraining BERT on Social Media (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that domain-specific BERT models can be improved when in-domain data is used for pretraining.
Approach: They propose to use Twitter and forum text as pretraining sources for two BERT models and use similarity measures to nominate in-domain data for pretraining.
Outcome: The proposed method can be used to improve performance on downstream tasks by using in-domain data.
Deciphering the Impact of Pretraining Data on Large Language Models through Machine Unlearning (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies have suggested that the composition of the pretraining corpus exerts a significant impact upon the performance of LLMs.
Approach: They analyze the impact of 48 datasets from 5 major categories of pretraining data of Large Language Models and measure their impacts on LLMs using benchmarks about nine major categories.
Outcome: The proposed analysis provides insights into the organization of data to support more efficient pretraining of Large Language Models.
Contextual Embeddings: When Are They Worth It? (2020.acl-main)

Copied to clipboard

Challenge: In recent years, rich contextual embeddings have enabled rapid progress on benchmarks like GLUE, but require significant computational resources during pretraining and during downstream task training and inference.
Approach: They empirically compare contextual embeddings with classic pretrained embedders and a random word embeddable with a simple baseline.
Outcome: The proposed models perform within 5 to 10% accuracy on industry-scale data.
Is BERT a Cross-Disciplinary Knowledge Learner? A Surprising Finding of Pre-trained Models’ Transferability (2021.findings-emnlp)

Copied to clipboard

Challenge: Using pre-trained language models, we can apply them to specialized domains such as scientific articles or clinical data.
Approach: They propose to pre-train BERT models on large text corpora and use them to generalize to token sequence classification applications.
Outcome: The models pre-trained on text classification tasks perform better than the models using task-specific knowledge and share non-trivial similarities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations