Papers by Bumjin Park
Incomplete Prompt Jailbreaks in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. |
| Approach: | They formalize incomplete prompt jailbreaks as incomplete prompts elicit harmful continuations . they identify two functional neurons that delay refusal until sentence termination . |
| Outcome: | The proposed model fails to generalize across content domains and attractor types . the proposed model can be used to perform more precise and robust IPJ defenses . |
Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly engaging in moral and ethical reasoning, where criteria for judgment are often unclear, even for humans. |
| Approach: | They propose a judgment strategy that integrates few-shot examples with reasoning prompts to mitigate this bias. |
| Outcome: | The proposed judgment strategy integrates few-shot examples with reasoning prompts to mitigate this bias. |