Papers with BEAST

1 papers
Towards Understanding the Robustness of Sparse Autoencoders (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are vulnerable to optimization-based jailbreak attacks that exploit internal gradient structure.
Approach: They propose to integrate pretrained Sparse Autoencoders into transformer residual streams at inference time without modifying model weights or blocking gradients.
Outcome: The proposed model reduces jailbreak success rate by 5x compared to baseline models . compared with models with weak white-box attacks, the proposed model is more robust .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations