Augmented Adversarial Trigger Learning (2025.findings-naacl)

Copied to clipboard

Challenge: Gradient optimization-based adversarial attack methods can generate jailbreak prompts or leak system prompts.
Approach: They propose an algorithm that enhances negative log-likelihood loss and augments it with auxiliary loss.
Outcome: The proposed approach outperforms current state-of-the-art techniques in nearly 100% of attacks while requiring 80% fewer queries.

Similar Papers

Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities (2025.naacl-long)

Copied to clipboard

Challenge: Recent research shows that Large Language Models (LLMs) are vulnerable to automated jailbreak attacks.
Approach: They propose a framework that crafts adversarial LLMs with enhanced jailbreak ability.
Outcome: ADV-LLM significantly reduces the computational cost of generating adversarial suffixes while achieving nearly 100% ASR on various open-source LLMs.
Universal Adversarial Triggers for Attacking and Analyzing NLP (D19-1)

Copied to clipboard

Challenge: Using adversarial triggers, a model can produce a specific prediction . adversarial attacks are useful for evaluation and interpretation .
Approach: They propose a gradient-guided search over tokens that finds short adversarial triggers that successfully trigger the target prediction.
Outcome: The proposed algorithm finds short trigger sequences that successfully trigger the target prediction.
Towards Open Domain Event Trigger Identification using Adversarial Domain Adaptation (2020.acl-main)

Copied to clipboard

Challenge: supervised event trigger identification models can generalize better across domains . prior work focused on annotating specific categories of events or narratives from specific domains.
Approach: They propose to use adversarial domain adaptation framework to build supervised event trigger identification models which can generalize better across domains.
Outcome: The proposed model improves on literature and news domains with no labeled data.
Prompt Optimization via Adversarial In-Context Learning (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to optimize prompts for in-context learning are based on adversarial learning and are computationally efficient and extensible to other LLMs and tasks.
Approach: They propose a method to optimize prompts for in-context learning by a generator and a discriminator.
Outcome: The proposed method improves state-of-the-art prompt optimization techniques on 13 generation and classification tasks including summarization, arithmetic reasoning, machine translation, data-to-text generation, and the MMLU and big-bench hard benchmarks.
AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt (2025.emnlp-main)

Copied to clipboard

Challenge: Existing RAG attacks rely on manipulating user queries, but exploit instructional prompts to manipulate RAG outputs covertly.
Approach: They propose an attack that exploits adversarial instructional prompts to manipulate RAG outputs . they propose a query generation strategy that simulates realistic linguistic variation in user queries .
Outcome: The proposed attack exploits instructional prompts to manipulate RAG outputs . it achieves up to 95.23% attack success rate while maintaining benign functionality .
Monte Carlo Tree Search Based Prompt Autogeneration for Jailbreak Attacks against LLMs (2025.coling-main)

Copied to clipboard

Challenge: Jailbreak attacks craft specific prompts or append adversarial suffixes to prompts, thereby inducing language models to generate harmful or unethical content and bypassing the model’s safety guardrails.
Approach: They propose a Monte Carlo Tree Search (MCTS) based Prompt Auto-generation (MPA) method to generate adversarial suffixes for valid jailbreak attacks.
Outcome: The proposed method outperforms existing methods on open-source and closed-source models and shows that it can generate harmful responses.
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods to defend against jailbreak attacks exploit vulnerabilities to elicit unintended or harmful outputs.
Approach: They propose a method to defend against jailbreak attacks by patching specific layers within large language models through self-augmented datasets.
Outcome: The proposed approach reduces harmfulness and attack success rate of jailbreak attacks without compromising utility for benign queries compared to previous methods.
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs (2025.acl-long)

Copied to clipboard

Challenge: cipher decoding, riddles, code execution embedded into model prompts bypass safety safeguards of large language models (LLMs) .
Approach: They introduce a novel class of adversarial jailbreak adversarials on large language models, termed Task-in-Prompt (TIP) attacks.
Outcome: The proposed techniques circumvent safeguards in six state-of-the-art language models, including GPT-4o and LLaMA 3.2, and consistently generate restricted content .
Universal Adversarial Attacks with Natural Triggers for Text Classification (2021.naacl-main)

Copied to clipboard

Challenge: Recent work has demonstrated the vulnerability of modern text classifiers to universal adversarial attacks, which are input-agnostic sequences of words added to text processed by classifier.
Approach: They propose a gradient-based search that aims to maximize the downstream classifier’s prediction loss by using an adversarially regularized autoencoder to generate triggers and propose heuristics to spot such attacks.
Outcome: The proposed algorithms reduce model accuracy while being less identifiable than prior models as per automatic detection metrics and human-subject studies.
A Simple and Efficient Learning-Style Prompting for LLM Jailbreaking (2026.findings-eacl)

Copied to clipboard

Challenge: Learning-style queries can reliably elicit harmful responses, highlighting a critical safety blind spot in modern LLMs.
Approach: They propose a new reframing paradigm that hides intention by learning from LLMs and uses 4 conceptual components to construct learning-style queries.
Outcome: The proposed framework achieves top attack success rates on most models and across malicious categories while maintaining high efficiency with concise prompts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations