Papers by Soomin Han
HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human–LLM Collaborative Writing (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models are increasingly used as coauthors in collaborative writing . however, this capability poses a serious safety risk . |
| Approach: | They propose a safety-utility balanced alignment approach to train LLMs to refuse harmful completions while remaining helpful on benign drafts. |
| Outcome: | The proposed method reduces harmful outputs without degrading performance on co-authoring capabilities. |