Papers by Jiayi Mao
“Not Aligned” is Not “Malicious”: Being Careful about Hallucinations of Large Language Models’ Jailbreak (2025.coling-main)
Copied to clipboard
| Challenge: | “Jailbreak” is a major safety concern of Large Language Models (LLMs). |
| Approach: | They propose a benchmarking framework to evaluate "jailbreak" outputs . they propose specialized validation framework to ensure outputs are useful malicious instructions . |
| Outcome: | The proposed framework enhances existing benchmarks to ensure outputs are useful . it also aims to evaluate the true potential of jailbroken outputs to cause harm to human society. |