Challenge: Spurious correlations are patterns that appear in datasets but do not represent genuine relationships.
Approach: They propose a more general form of counterfactual data augmentation that tackles multiple biases . they propose 'CoBA' that decomposes text into subject-predicate-object triples and modifies them to disrupt spurious correlations.
Outcome: The proposed framework reduces biases and strengthens out-of-distribution resilience.

Similar Papers

Decorrelate Irrelevant, Purify Relevant: Overcome Textual Spurious Correlations from a Feature Perspective (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to debiase samples with biased features obstructs the model in learning from non-biased parts of the samples.
Approach: They propose to eliminate spurious correlations in a fine-grained manner from a feature space perspective by using Random Fourier Features and weighted re-sampling to decorrelate dependencies between features.
Outcome: The proposed method eliminates spurious correlations in a fine-grained manner from a feature space perspective.
An Investigation of the (In)effectiveness of Counterfactually Augmented Data (2022.acl-long)

Copied to clipboard

Challenge: Pretrained language models tend to rely on spurious correlations and generalize poorly to out-of-distribution (OOD) data.
Approach: They propose to use counterfactually-augmented data (CAD) to identify robust features that are invariant under distribution shift to train models for OOD generalization.
Outcome: The proposed model can learn robust features that are invariant under distribution shifts, but lacks spurious correlations, and may exacerbate existing correlations.
Consistent Document-level Relation Extraction via Counterfactuals (2024.findings-emnlp)

Copied to clipboard

Challenge: Document-level relation extraction models trained on factual data exhibit inconsistent behavior, relying on spurious signals such as specific entities and external knowledge to extract triples.
Approach: They propose a counterfactual data generation approach for document-level relation extraction datasets using entity replacement to generate triples from factual data.
Outcome: The proposed approach extracts triples from factual data but fails on counterfactual modification.
Counterfactual Inference for Text Classification Debiasing (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to capture unintended dataset biases are expensive and require elaborate balancing strategies.
Approach: They propose a model-agnostic text classification debiasing framework which can effectively avoid employing data manipulations or designing balancing mechanisms.
Outcome: The proposed framework can effectively avoid data manipulations or designing balancing mechanisms to capture unintended dataset biases.
NeuroCounterfactuals: Beyond Minimal-Edit Counterfactuals for Richer Data Augmentation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to produce counterfactuals rely on small perturbations via minimal edits, resulting in simplistic changes.
Approach: They propose a novel approach to produce counterfactuals that allow for larger edits and linguistic diversity while still bearing similarity to the original document.
Outcome: The proposed approach outperforms existing methods for generalizing natural language models under select settings.
Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets (2022.acl-long)

Copied to clipboard

Challenge: Natural language processing models exploit spurious correlations between features and labels in datasets to perform well only within the distributions they are trained on.
Approach: They propose to generate a debiased version of a dataset and replace it with training data to train a model that is generalised to different task distributions.
Outcome: The proposed method outperforms or performs comparable to state-of-the-art debiasing strategies on a large suite of debiased, out-of distribution, and adversarial test sets.
Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity Typing (2022.emnlp-main)

Copied to clipboard

Challenge: Existing entity typing models are subject to spurious correlations due to shortcuts and biased training.
Approach: They propose a method to augment existing model biases by combining spurious correlations with debiasedcounterparts to improve generalization.
Outcome: The proposed method improves generalization of different entity typing models on the original and debiased test sets.
Can We Improve Model Robustness through Secondary Attribute Counterfactuals? (2021.emnlp-main)

Copied to clipboard

Challenge: Recent research has explored how models rely on spurious correlations and how counterfactual data augmentation (CDA) can mitigate such issues.
Approach: They propose a context-aware methodology which takes into account the impact of secondary attributes on the model’s predictions and increases sensitivity for secondary attributes over reweighted counterfactually augmented data.
Outcome: The proposed approach improves sliced accuracy on the original dataset by 7% compared to existing methods and provides guidelines to extend this to other tasks.
Counterfactual Augmentation for Multimodal Learning Under Presentation Bias (2023.findings-emnlp)

Copied to clipboard

Challenge: In real-world machine learning systems, labels are often derived from user behaviors that the system wishes to encourage.
Approach: They propose a method for correcting presentation bias using generated counterfactual labels by augmentation of the labels by the user.
Outcome: The proposed method improves performance in an oracle setting compared to uncorrected models and existing bias-correction methods.
Counterfactual Debiasing for Fact Verification (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for debiasing factchecking models learn such biases instead of understanding the semantic relationship between the claim and evidence.
Approach: They propose a counterfactual framework CLEVER which is augmentation-free and mitigates biases on the inference stage.
Outcome: The proposed method is augmentation-free and mitigates biases on the inference stage.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations