Papers by Yu-Ling Hsu
Jailbreaking with Universal Multi-Prompts (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have seen rapid development in recent years, but ethical concerns and new types of attacks have emerged. |
| Approach: | They propose a prompt-based method to jailbreak large language models using universal multi-prompts and an approach for defense that outperforms existing techniques. |
| Outcome: | The proposed method outperforms existing techniques for jailbreaking LLMs using universal multi-prompts. |