Papers by Maithili Joshi
SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) with safe-alignment training are vulnerable to jailbreak attacks, causing malicious users to generate harmful outputs. |
| Approach: | They propose a safe-alignment jailbreak method that bypasses the middle-to-late layers of large language models by a residual connection. |
| Outcome: | The proposed method improves by 51% over the best performing baseline GCG on HarmBench test set. |