Challenge: Existing methods for steering concept vectors suffer from noisy features in diverse datasets that undermine steering robustness.
Approach: They propose a Sparse Autoencoder-Denoised Concept Vector (SDCV) which selectively keeps the most discriminative SAE latents while reconstructing hidden representations.
Outcome: The proposed method improves steering success rates by 4-16% across six challenging concepts while maintaining topic relevance.

Similar Papers

SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities in natural language understanding and generation, but controlling their behavior remains a challenge.
Approach: They propose a supervised steering approach that operates in sparse, interpretable representation spaces.
Outcome: The proposed approach achieves higher success rates with minimal degradation in generation quality compared to existing methods.
Toward Efficient Sparse Autoencoder-Guided Steering for Improved In-Context Learning in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Sparse autoencoders (SAEs) have emerged as a powerful analytical tool in mechanistic interpretability for large language models (LLMs).
Approach: They propose a novel approach that leverages SAEs to enhance the general in-context learning performance of large language models (LLMs).
Outcome: The proposed method yields a 3.5% improvement across diverse text classification tasks and exhibits greater robustness to hyperparameter variations compared to standard steering approaches.
Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit impressive abilities in various domains such as text generation, instruction following, and reasoning.
Approach: They propose a method to decompose the activations of Large Language Models into a sparse linear combination of SAE features.
Outcome: The proposed method shows that some features are strongly related to specific languages, while others are unaffected by ablating them.
CRISP: Persistent Concept Unlearning via Sparse Autoencoders (2026.acl-long)

Copied to clipboard

Challenge: Recent work has explored sparse autoencoders (SAEs) to perform precise interventions on monosemantic features, but most SAE-based methods operate at inference time, which does not create persistent changes in the model’s parameters.
Approach: They propose a parameter-efficient method for persistent concept unlearning using SAEs that automatically identifies salient SAE features across multiple layers and suppresses their activations.
Outcome: The proposed method outperforms previous methods on safety-critical unlearning tasks from the WMDP benchmark, successfully removing harmful knowledge while preserving general and in-domain capabilities.
Deciphering Cultural Representations in Large Language Models via Sparse Autoencoders (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has identified so-called cultural neurons, but individual neurons are often polysemous, conflating abstract cultural knowledge with surface-level lexical cues due to superposition.
Approach: They apply Sparse Autoencoders to decompose LLM activations into sparse, interpretable feature representations that disentangle culturally selective features.
Outcome: The proposed model disentangles culturally selective features from paraphrasing and task formats, indicating abstraction beyond lexical correlations.
Unsupervised Concept Vector Extraction for Bias Control in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases.
Approach: They propose a method that extracts concept representations via probability weighting without labeled data and efficiently selects a steering vector for measuring and manipulating the model’s representation.
Outcome: The proposed method can be used to predict gender bias and generalizes to racial bias.
SAEs Are Good for Steering – If You Select the Right Features (2025.emnlp-main)

Copied to clipboard

Challenge: Sparse Autoencoders (SAEs) can learn a decomposition of a model’s latent space by analyzing the input tokens that activate them.
Approach: They propose an unsupervised approach to learn a decomposition of a model’s latent space by analyzing the input tokens that activate them.
Outcome: The proposed approach matches the performance of existing supervised methods by identifying features with low output scores and identifying them with input and output scores.
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering (2026.acl-long)

Copied to clipboard

Challenge: Existing methods struggle to control fine-grained reasoning strategies due to conceptual entanglement in LRMs’ hidden states.
Approach: They propose to decompose strategy-entangled hidden states into a disentangled feature space by using Sparse Autoencoders to identify the few strategy-specific features from the vast pool of SAE features.
Outcome: The proposed method outperforms existing methods by 15% in control effectiveness.
Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders (2025.findings-acl)

Copied to clipboard

Challenge: Existing algorithms for AI text detection lack interpretability, limiting their reliability in highstakes applications.
Approach: They extend existing ATD frameworks by using Sparse Autoencoders to extract features from Gemma-2-2b residual stream.
Outcome: The proposed algorithms can extract human-interpretable features from Gemma-2-2b model.
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are ubiquitous in user-facing applications, yet they still generate undesirable toxic outputs, including profanity, vulgarity, and derogatory remarks.
Approach: They leverage sparse autoencoders to identify toxicity-related directions in residual stream of large language models and perform targeted activation steering using the corresponding decoder vectors.
Outcome: The proposed models surpass baselines in reducing toxicity by up to 20%, though fluency can degrade noticeably on GPT-2 Small and Gemma-2-2B.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations