EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks (D19-1)
Copied to clipboard
| Challenge: | Existing data augmentation techniques for text classification are difficult to implement and cost a high amount of money. |
| Approach: | They propose to use four simple but powerful operations to boost performance on text classification tasks to improve synonym replacement, random insertion, random swap, and random deletion. |
| Outcome: | The proposed techniques improve performance on five classification tasks and are particularly useful for smaller datasets. |
Similar Papers
AEDA: An Easier Data Augmentation Technique for Text Classification (2021.findings-emnlp)
Copied to clipboard
| Challenge: | AEDA is an easier data augmentation technique than EDA. |
| Approach: | They propose an augmentation technique that includes only random insertion of punctuation marks into the original text. |
| Outcome: | The proposed method is easier to implement for data augmentation than EDA method. |
EPiDA: An Easy Plug-in Data Augmentation Framework for High Performance Text Classification (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing methods for data augmentation do not fully exploit the potential of DA in NLP. |
| Approach: | They propose an easy and plug-in framework for data augmentation to support effective text classification. |
| Outcome: | The proposed framework outperforms existing methods in most cases, but not using agent networks or pre-trained generation networks. |
An Analysis of Simple Data Augmentation for Named Entity Recognition (2020.coling-main)
Copied to clipboard
| Challenge: | Recent studies have focused on using data augmentation techniques on sentence-level and sentence-pair natural language processing tasks such as text classification. |
| Approach: | They propose to use data augmentation techniques for named entity recognition to increase model performance. |
| Outcome: | The proposed techniques boost performance for both recurrent and transformer-based models, especially for small training sets. |
FlipDA: Effective and Robust Data Augmentation for Few-Shot Learning (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods for text data augmentation are limited to simple tasks and weak baselines. |
| Approach: | They propose a data augmentation method FlipDA that uses a generative model and a classifier to generate label-flipped data. |
| Outcome: | The proposed method improves many tasks while not negatively affecting the others. |
All You Need is Attention: Lightweight Attention-based Data Augmentation for Text Classification (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to augment text classification tasks require extensive dataset training. |
| Approach: | They propose a method that uses attention mechanisms to exchange semantically similar words between sentences to generate a greater diversity of synthetic sentences compared to simpler operations like random insertions. |
| Outcome: | The proposed method consistently outperforms baseline methods across diverse text classification conditions. |
Text Smoothing: Enhance Various Data Augmentation Methods on Text Classification Tasks (2022.acl-short)
Copied to clipboard
| Challenge: | Experimental results show text smoothing outperforms data augmentation methods by a substantial margin. |
| Approach: | They propose to use a masked language model to convert a token to a smoothed representation by converting a sentence from its one-hot representation to 'controllable smoothes' they propose to combine text smoothing with other data augmentation methods to achieve better performance. |
| Outcome: | The proposed method outperforms mainstream data augmentation methods by a substantial margin on different datasets in a low-resource regime. |
Boosting Text Augmentation via Hybrid Instance Filtering Framework (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing text augmentation methods generate instances with shifted feature spaces, which leads to a drop in performance on large datasets. |
| Approach: | They propose a hybrid instance-filtering framework that generates instances with shifted feature spaces, which leads to a drop in performance on augmented data. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on three classification tasks and nine public datasets. |
SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation (D18-1)
Copied to clipboard
| Challenge: | Existing methods for data augmentation for text-based tasks such as machine translation are limited due to noise and noise. |
| Approach: | They propose a data augmentation policy with desirable properties as an optimization problem and propose 'SwitchOut' switchout randomly replaces words in both the source and target sentences with other random words from their corresponding vocabularies. |
| Outcome: | The proposed method outperforms strong alternatives such as word dropout on three translation datasets. |
On Evaluation Protocols for Data Augmentation in a Limited Data Scenario (2025.coling-main)
Copied to clipboard
| Challenge: | Textual data augmentation (DA) is a prolific field of study where novel techniques to create artificial data are regularly proposed. |
| Approach: | They propose to use textual data augmentation (DA) to generate new sentences for text classification in a limited data setting. |
| Outcome: | The proposed methods perform better on small data settings and on large datasets, but they are not as effective on large data sets. |
AutoAugment Is What You Need: Enhancing Rule-based Augmentation Methods in Low-resource Regimes (2024.eacl-srw)
Copied to clipboard
| Challenge: | Existing methods for text data augmentation suffer from potential semantic damage due to the discrete nature of sentences. |
| Approach: | They propose to adapt AutoAugment to solve this problem by using softEDA to increase text data. |
| Outcome: | The proposed method can boost existing augmentation methods and enhance cutting-edge pretrained language models. |