Papers by Sibei Yang
DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing safety mechanisms for Large Language Models (LLMs) are inadequate to protect against jailbreak attacks, resulting in performance degradation on general tasks. |
| Approach: | They propose a method that directly updates a minimal set of relevant parameters to neutralize harmful behaviors while preserving the model’s utility. |
| Outcome: | The proposed model outperforms baseline methods in mitigating jailbreak attacks while preserving the model’s utility. |
Don’t Say No: Jailbreaking LLM by Suppressing Refusal (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are vulnerable to "jailbreaking" attacks where crafted prompts manipulate them into producing toxic content. |
| Approach: | They propose to improve the target loss objective by combining a cosine decay schedule method with refusal suppression to achieve higher success rates. |
| Outcome: | The proposed approach outperforms baseline attacks and achieves state-of-the-art attack success rates. |