Papers by Elena Golimblevskaia
FADE: Why Bad Descriptions Happen to Good Features (2025.findings-acl)
Copied to clipboard
Bruno Puri, Aakriti Jain, Elena Golimblevskaia, Patrick Kahardipraja, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin
| Challenge: | Recent advances in mechanistic interpretability have highlighted the potential of automating interpretability pipelines in analyzing the latent representations within LLMs. |
| Approach: | They propose a framework for automatically evaluating feature-to-description alignment that measures alignment across four key metrics and quantifies the causes of misalignment. |
| Outcome: | The proposed framework evaluates alignment across four key metrics and quantifies the causes of misalignment between features and descriptions. |