Papers by Haon Park
sudo rm -rf agentic_security (2025.acl-industry)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used as computer-use agents . authors present a novel attack framework that bypasses refusal-trained safeguards . |
| Approach: | They propose a new attack framework that bypasses refusal-trained safeguards in LLMs . SUDO iteratively refines its attacks based on a built-in refusal feedback . authors highlight need for robust, context-aware safeguards if LLM is to be used . |
| Outcome: | The proposed framework bypasses refusal-trained safeguards in commercial agents . it achieves a stark attack success rate of 24.41% (with no refinement) and up to 41.33% (by iterative refinement). |
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs (2026.acl-long)
Copied to clipboard
Dasol Choi, DongGeon Lee, Brigitta Jesica Kartono, Helena Berndt, Taeyoun Kwon, Joonwon Jang, Haon Park, Hwanjo Yu, Minsuk Kahng
| Challenge: | Large language models are being rapidly adopted across a wide range of domains, including healthcare, finance, and the public sector. |
| Approach: | They propose a framework to evaluate whether large language models comply with policies . they apply COMPASS to eight diverse industry scenarios to validate models . |
| Outcome: | The proposed framework evaluates whether LLMs comply with allowlist and denylist policies. |
One-Shot is Enough: Consolidating Multi-Turn Attacks into Efficient Single-Turn Prompts for LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | a novel framework for consolidating multi-turn adversarial “jailbreak” prompts into single-turn queries is presented in a journal of computational linguistics. |
| Approach: | They propose a framework for consolidating adversarial “jailbreak” prompts into single-turn queries. |
| Outcome: | The proposed framework outperforms the original multi-turn attacks by up to 17.5 % in absolute ASR . it reduces token usage by more than half on average, and provides a powerful tool for large-scale red-teaming . |