Papers by Sebastian Lapuschkin
FADE: Why Bad Descriptions Happen to Good Features (2025.findings-acl)
Copied to clipboard
Bruno Puri, Aakriti Jain, Elena Golimblevskaia, Patrick Kahardipraja, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin
| Challenge: | Recent advances in mechanistic interpretability have highlighted the potential of automating interpretability pipelines in analyzing the latent representations within LLMs. |
| Approach: | They propose a framework for automatically evaluating feature-to-description alignment that measures alignment across four key metrics and quantifies the causes of misalignment. |
| Outcome: | The proposed framework evaluates alignment across four key metrics and quantifies the causes of misalignment between features and descriptions. |
From Weights to Activations: Is Steering the Next Frontier of Adaptation? (2026.acl-long)
Copied to clipboard
Simon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich, Sebastian Lapuschkin, Wojciech Samek, Vera Schmitt
| Challenge: | Pre-trained large language models are the basis of a wide range of NLP tasks. |
| Approach: | They propose to use parameter updates and parameter-efficient adaptation to modify behavior of large language models. |
| Outcome: | The proposed method enables local and reversible behavioral change without parameter updates. |