Analyzing (In)Abilities of SAEs via Formal Languages (2025.naacl-long)

Copied to clipboard

Challenge: Autoencoders have been used for finding interpretable and disentangled features underlying neural network representations in both image and text domains, but there is a lack of corresponding results for the text domain.
Approach: They propose to train sparse autoencoders (SAEs) on a synthetic testbed of formal languages to find interpretable latents in models trained on formal languages.
Outcome: The proposed approach promotes learning of causally relevant features in a formal language setting.

Similar Papers

On the Versatility of Sparse Autoencoders for In-Context Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Sparse autoencoders (SAEs) are emerging as a key analytical tool in interpretability for large language models.
Approach: They propose to use SAEs to extract knowledge from billions of tokens for sparse reconstruction.
Outcome: The proposed model can extract knowledge from billions of tokens for sparse reconstruction.
Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders (2026.findings-eacl)

Copied to clipboard

Challenge: Concept-based explanations for large language models are not well understood in text classification.
Approach: They propose a model with a specialized classifier head and activation rate sparsity loss for sentence classification . they compare it to existing models with HI-Concept and ConceptShap .
Outcome: The proposed model improves both the causality and interpretability of the extracted features.
Sparse Autoencoder Features for Classifications and Transferability (2025.emnlp-main)

Copied to clipboard

Challenge: Sparse Autoencoders (SAEs) provide potential for uncovering structured, human-interpretable representations in Large Language Models (LLMs).
Approach: They analyze SAEs for interpretable feature extraction from Large Language Models in safety-critical classification tasks.
Outcome: The proposed framework outperforms hidden-state and BoW models while demonstrating cross-lingual toxicity detection and visual classification tasks.
Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders (2025.findings-acl)

Copied to clipboard

Challenge: Existing algorithms for AI text detection lack interpretability, limiting their reliability in highstakes applications.
Approach: They extend existing ATD frameworks by using Sparse Autoencoders to extract features from Gemma-2-2b residual stream.
Outcome: The proposed algorithms can extract human-interpretable features from Gemma-2-2b model.
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Datasets Dependency (2025.acl-srw)

Copied to clipboard

Challenge: Sparse Autoencoders (SAEs) have emerged as a promising solution for decomposing large language model representations into interpretable features.
Approach: They propose a method that trains SAEs on the model’s own synthetic dataset and a model-specific model to capture model-internal features.
Outcome: The proposed method outperforms SAEs trained on web-based datasets and exhibits lower Fake Feature Ratio in 5 out of 7 models.
SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities in natural language understanding and generation, but controlling their behavior remains a challenge.
Approach: They propose a supervised steering approach that operates in sparse, interpretable representation spaces.
Outcome: The proposed approach achieves higher success rates with minimal degradation in generation quality compared to existing methods.
Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit impressive abilities in various domains such as text generation, instruction following, and reasoning.
Approach: They propose a method to decompose the activations of Large Language Models into a sparse linear combination of SAE features.
Outcome: The proposed method shows that some features are strongly related to specific languages, while others are unaffected by ablating them.
AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders (2026.eacl-long)

Copied to clipboard

Challenge: Feature steering reduces Whisper’s false speech detections by 70% with negligible WER increase, demonstrating real-world applicability.
Approach: They train Sparse Autoencoders across all encoder layers of Whisper and HuBERT and evaluate their stability, interpretability, and practical utility.
Outcome: The proposed models capture general acoustic and semantic information as well as specific events, including environmental noises and paralinguistic sounds, and disentangle them effectively.
SAEs Are Good for Steering – If You Select the Right Features (2025.emnlp-main)

Copied to clipboard

Challenge: Sparse Autoencoders (SAEs) can learn a decomposition of a model’s latent space by analyzing the input tokens that activate them.
Approach: They propose an unsupervised approach to learn a decomposition of a model’s latent space by analyzing the input tokens that activate them.
Outcome: The proposed approach matches the performance of existing supervised methods by identifying features with low output scores and identifying them with input and output scores.
Constructing Interpretable Features from Compositional Neuron Groups (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for analyzing LLMs rely on dictionary learning with sparse autoencoders (SAEs) however, SAEs struggle in causal evaluations and lack intrinsic interpretability, as their learning is not explicitly tied to the computations of the model.
Approach: They propose to decompose MLP activations with semi-nonnegative matrix factorization (SNMF) such that the learned features are mapped to their activating inputs, making them directly interpretable.
Outcome: Experiments on Llama 3.1, Gemma 2 and GPT-2 show that SNMF derived features outperform SAEs and a strong supervised baseline on causal steering while aligning with human-interpretable concepts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations