Papers with BEAST
Towards Understanding the Robustness of Sparse Autoencoders (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are vulnerable to optimization-based jailbreak attacks that exploit internal gradient structure. |
| Approach: | They propose to integrate pretrained Sparse Autoencoders into transformer residual streams at inference time without modifying model weights or blocking gradients. |
| Outcome: | The proposed model reduces jailbreak success rate by 5x compared to baseline models . compared with models with weak white-box attacks, the proposed model is more robust . |