Papers by Jinyuan Jia
Jailbreak Open-Sourced Large Language Models via Enforced Decoding (2024.acl-long)
Copied to clipboard
Hangfan Zhang, Zhimeng Guo, Huaisheng Zhu, Bochuan Cao, Lu Lin, Jinyuan Jia, Jinghui Chen, Dinghao Wu
| Challenge: | Existing studies show that Large Language Models can be misused to generate undesired content. |
| Approach: | They propose to use large language models to manipulate the generation process to generate undesired content without heavy computations or prompt designs. |
| Outcome: | The proposed method shows that open-sourced large language models could be misused to generate undesired content without heavy computations or prompt designs. |
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding (2024.acl-long)
Copied to clipboard
| Challenge: | Despite advances in large language models, they face substantial challenges in terms of safety. |
| Approach: | They develop a safety-aware decoding strategy for large language models to defend against jailbreak attacks. |
| Outcome: | The proposed strategy outperforms six defense methods against jailbreak attacks on five LLMs. |
Foot-In-The-Door: A Multi-turn Jailbreak for LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly integrated into real-world applications, requiring a high level of safety and alignment. |
| Approach: | They propose a multi-turn jailbreak method that leverages foot-in-the-door principles to escalate malicious intent of user queries through intermediate bridge prompts and aligns the model’s response by itself to induce toxic responses. |
| Outcome: | The proposed method achieves an average attack success rate of 94% across seven widely used models outperforming existing state-of-the-art methods. |
PIArena: A Platform for Prompt Injection Evaluation (2026.acl-long)
Copied to clipboard
| Challenge: | OWASP identifies prompt injection as the top-1 security risk for large language models (LLMs). |
| Approach: | They propose a unified platform for prompt injection evaluation that integrates state-of-the-art attacks and defenses into a platform. |
| Outcome: | The proposed attack exploits state-of-the-art defenses and generalizes them on diverse datasets and attacks. |