Challenge: a study of refusal in instruction-tuned language models identifies latent features that causally mediate refusal behaviors.
Approach: They conduct a mechanistic study of refusal in instruction-tuned LLMs using sparse autoencoders . they identify latent features that causally mediate refusal behaviors using sparsed autoencoding .
Outcome: The proposed method validates refusal-related features across multiple datasets.

Similar Papers

Feature Drift: How Fine-Tuning Repurposes Representations in LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Sparse autoencoders (SAEs) are a powerful tool for interpreting neural networks by extracting concepts (features) represented in their activations.
Approach: They propose to use Sparse Autoencoders to extract concepts from their activations to explain how fine-tuning changes model capabilities.
Outcome: The proposed model recombines existing concepts rather than learning new ones, and shows that it is a better explanation for how fine-tuning changes model capabilities.
Uncovering Sentiment Analysis Circuit in Large Language Model (2026.acl-long)

Copied to clipboard

Challenge: Prior work has shown that sentiment is encoded linearly in LLM representations, but their ability to utilize this information remains fragile to prompt variations.
Approach: They propose a simple inference-time intervention method that amplifies circuit features to compensate for insufficient activation.
Outcome: The proposed method improves on a sentiment analysis circuit with sparse autoencoders and circuit-level analysis.
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Sparse Autoencoders (SAEs) can disentangle complex features into more interpretable components.
Approach: They propose to use Sparse Autoencoders to disentangle LLM features into more interpretable components.
Outcome: The proposed method disentangles complex features into more interpretable components.
Sparse Latents Steer Retrieval-Augmented Generation (2025.acl-long)

Copied to clipboard

Challenge: In this study, we uncover interpretable latents that govern RAG behavior in large language models . Sparse Autoencoders are used to control large language model (LLM) behavior .
Approach: They leverage Sparse Autoencoders within the LLaMA Scope to uncover latents that govern RAG behaviors.
Outcome: The proposed model can be used to control large language models without architectural modifications.
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries (2026.findings-acl)

Copied to clipboard

Challenge: Existing refusal strategies for unanswerable and underspecified user queries are brittle due to model hallucinations or add complexity and overhead.
Approach: They propose a latent-signal refusal mechanism that predicts query answerability from hidden activations of an LLM.
Outcome: The proposed scheme reduces schema noise and sparse, localized question–schema mismatch cues that indicate unanswerability.
Incomplete Prompt Jailbreaks in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests.
Approach: They formalize incomplete prompt jailbreaks as incomplete prompts elicit harmful continuations . they identify two functional neurons that delay refusal until sentence termination .
Outcome: The proposed model fails to generalize across content domains and attractor types . the proposed model can be used to perform more precise and robust IPJ defenses .
Deciphering Cultural Representations in Large Language Models via Sparse Autoencoders (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has identified so-called cultural neurons, but individual neurons are often polysemous, conflating abstract cultural knowledge with surface-level lexical cues due to superposition.
Approach: They apply Sparse Autoencoders to decompose LLM activations into sparse, interpretable feature representations that disentangle culturally selective features.
Outcome: The proposed model disentangles culturally selective features from paraphrasing and task formats, indicating abstraction beyond lexical correlations.
The Unintended Trade-off of AI Alignment: Balancing Hallucination Mitigation and Safety in LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Hallucination in large language models has been studied, but a side effect remains unrecognized . a new study examines the trade-off between truthfulness and safety alignment .
Approach: They propose a method that disentangles hallucination from hallucinian features using sparse autoencoders.
Outcome: The proposed method preserves refusal behavior and task utility while maintaining safety alignment.
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders (2025.emnlp-main)

Copied to clipboard

Challenge: Sparse Autoencoders (SAEs) have recently emerged as powerful tools for interpreting and steering the internal representations of large language models (LLMs).
Approach: They propose a method that identifies the most influential latents by incorporating output-side gradient information.
Outcome: The proposed method identifies the most influential latents by incorporating output-side gradient information.
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective (2025.findings-acl)

Copied to clipboard

Challenge: Current approaches to value alignment focus on a few core values, such as helpfulness, harmlessness, and honesty.
Approach: They propose to use latent causal value graphs to guide two lightweight value-steering methods . role-based prompting and sparse autoencoder (SAE) steering are also used .
Outcome: Experiments on Gemma-2B-IT and Llama3-8B- IT show that the proposed methods are effective and controllable.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations