Challenge: Backdoor attacks manipulate model predictions by inserting malicious "poison" instances that contain a specific pattern or "trigger."
Approach: They propose an attack that inserts style-based triggers into training and test data by using a poison selection technique to improve the effectiveness of both LLMBkd and existing backdoor attacks.
Outcome: The proposed attack achieves high success rates across a wide range of styles with little effort and no model training.

Similar Papers

Triggerless Backdoor Attack for NLP Tasks with Clean Labels (2022.naacl-main)

Copied to clipboard

Challenge: Backdoor attacks are a new threat to neural natural language processing models due to the fragility and lack of interpretability of NLP models.
Approach: They propose a method to perform backdoor attacks without an external trigger . they propose to use clean-labeled examples to generate poisoned clean-labelled examples .
Outcome: The proposed strategy is effective and hard to defend due to its triggerless nature.
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning (2024.emnlp-main)

Copied to clipboard

Challenge: In-context learning has shown high efficacy in several NLP tasks, especially in few-shot settings.
Approach: They propose a backdoor attack method that poisons demonstration examples and poisons the demonstration context, preserving the model's generality.
Outcome: The proposed method can make models behave in alignment with predefined intentions without fine-tuning the model.
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that Large Language Models (LLMs) are susceptible to backdoor attacks, where triggers embedded in poisoned data can maliciously alter LLMs’ behaviors.
Approach: They propose to leverage LLMs' generative capabilities to generate human-readable explanations for their decisions, enabling direct comparisons between explanations of clean and poisoned data.
Outcome: The proposed model produces coherent explanations for clean inputs but logically flawed explanations on poisoned data.
PKAD: Pretrained Knowledge is All You Need to Detect and Mitigate Textual Backdoor Attacks (2024.findings-emnlp)

Copied to clipboard

Challenge: Current defense methods can be classified into inference-time and training-time ones based on their execution phase.
Approach: They propose a two-stage poison detection strategy using pre-trained language models to detect poisoned samples before model training.
Outcome: The proposed method achieves better performance than current methods more quickly and with fewer training costs.
Backdoor NLP Models via AI-Generated Text (2024.lrec-main)

Copied to clipboard

Challenge: Existing attacks disregard fluency and semantic fidelity of poisoned text, rendering it easily detectable.
Approach: They propose to use AI-generated poisoned text to attack NLP models by establishing covert associations between trigger patterns and target labels without affecting normal accuracy.
Outcome: The proposed method achieves effective attacks while maintaining fluency and semantic similarity across all scenarios.
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Generative large language models (LLMs) have remarkable performance in generation tasks, but datasets used to train or fine-tune these models are often not disclosed to users.
Approach: They develop an inference time defense called CleanGen to mitigate backdoor attacks for generation tasks in large language models.
Outcome: The proposed inference time defense achieves lower attack success rates (ASR) compared to baseline defenses for all five backdoor attacks.
Rethinking Backdoor Detection Evaluation for Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing backdoor detection methods have high accuracy in detecting backdoored models, but they are not robust enough to detect backdoors in the wild.
Approach: They examine the robustness of backdoor detectors by manipulating different factors during backdoor planting.
Outcome: The proposed methods are able to detect backdoors in the wild, but they lack robustness against backdoor attacks.
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: ProAttack is a novel and efficient method for performing clean-label backdoor attacks based on the prompt, which uses the prompt itself as a trigger.
Approach: They propose a method for performing clean-label backdoor attacks based on the prompt, which uses the prompt itself as a trigger.
Outcome: The proposed method achieves state-of-the-art performance on several NLP tasks, particularly in few-shot settings.
BFClass: A Backdoor-free Text Classification Framework (2021.findings-emnlp)

Copied to clipboard

Challenge: Various trigger design strategies have been explored to attack text classifiers, however, defending such attacks remains an open problem.
Approach: They propose a backdoor-free training framework that poisons a subset of training data by injecting trigger patterns and setting their labels as the target labels.
Outcome: The proposed framework can detect all the triggers, remove 95% of poisoned training samples with very limited false alarms, and achieve almost the same performance as the models trained on benign training data.
Textual Backdoor Attacks Can Be More Harmful via Two Simple Tricks (2022.emnlp-main)

Copied to clipboard

Challenge: Existing textual backdoor attacks are vulnerable to backdoors . researchers add extra training task to distinguish poisoned and clean data .
Approach: They propose two tricks that make existing backdoor attacks much more harmful . first trick is to add an extra task to distinguish poisoned and clean data . second trick is using all the clean training data rather than the original clean data.
Outcome: The proposed tricks can significantly improve attack performance in three tough situations including clean data fine-tuning, low-poisoning-rate, and label-consistent attacks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations