Papers by Hiroki Furuta
Understanding Emergent Misalignment via Feature Superposition Geometry (2026.acl-long)
Copied to clipboard
| Challenge: | Emergent misalignment is a problem for large language models (LLMs) fine-tuning on narrow tasks can induce harmful behaviors despite no explicit supervision. |
| Approach: | They propose a mechanistic account based on the geometry of feature superposition . they propose to use sparse autoencoders to identify misalignment-inducing features . |
| Outcome: | The proposed model outperforms random removal and stronger mitigations than LLM-as-a-judge filtering. |