Papers with Superposition
SAFR: Neuron Redistribution for Interpretability (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies on controlling neuron distribution for interpretability have focused on focusing on monosemanticity instead of focusing solely on feature interactions. |
| Approach: | They propose a method to regularize feature superposition by encoding representations of multiple features within a single neuron. |
| Outcome: | The proposed method improves model interpretability without compromising prediction performance. |