Papers by Aina Sui
RAMP: Risk-Aware Multi-Turn Planning for Jailbreak Red-Teaming (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for evaluating the safety of large language models rely on heuristic strategies or trained attack agents. |
| Approach: | They propose a method for multi-turn jailbreaking that iteratively plans and executes each turn via a Judge, a Transitioner, and a Planner. |
| Outcome: | The proposed framework achieves strong attack performance across open-source and closed-source target models while remaining effective under stricter turn budgets. |