Papers by Amit LeVi
Jailbreak Attack Initializations as Extractors of Compliance Directions (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Safety-aligned LLMs respond to prompts with compliance or refusal, each corresponding to distinct directions in the model’s activation space. |
| Approach: | They propose an initialization framework that aims to project unseen prompts further along compliance directions. |
| Outcome: | The proposed initialization framework achieves an increased attack success rate and reduced computational overhead, highlighting the fragility of safety-aligned LLMs. |