Papers by Michael Karasik
In-Context Representation Hijacking (2026.acl-long)
Copied to clipboard
| Challenge: | In-context representation hijacking attacks on large language models are insufficient to prevent harm, yet the mechanisms underlying this behavior remain unclear. |
| Approach: | They propose a simple in-context representation hijacking attack that replaces a harmful keyword with a benign token, and embeds the harmful semantics under a euphemism. |
| Outcome: | The proposed attack is optimized for open-source and achieves strong success rates on closed-source systems, reaching 74% on Llama-3.3-70B-Instruct with a single-sentence context override. |