Papers by Runyi Hu
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing jailbreak methods struggle to balance effectiveness with robustness against adaptive safety mechanisms. |
| Approach: | They propose a novel approach that targets Large Reasoning Models through an adaptive encryption pipeline designed to overwhelm their reasoning capabilities. |
| Outcome: | The proposed approach achieves an attack success rate of 85.6% on OpenAI GPT-o4-mini, outperforming state-of-the-art baselines by a significant margin of 17.2%. |