Challenge: Existing methods for inference-time steering fail to be effective, utility-preserving and training-efficient due to rigid, one-size-fits-all designs and limited adaptability.
Approach: They propose a steering framework that decomposes inference-time steering into two stages . they propose 'conditional steering' mechanism that preserves model utility by avoiding unnecessary steering . a 'mixture-of-Steering-Experts' mechanism captures multimodal nature of desired steering behaviors .
Outcome: The proposed framework outperforms the state-of-the-art methods on safety and truthfulness benchmarks.

Similar Papers

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for fine-tuning large language models often ignore token-level causal influence and underutilize model logits.
Approach: They propose a novel approach that uses a gradient-based approach to identify influential tokens and construct directional steering vectors based on their contribution to preferred over dispreferred outputs.
Outcome: The proposed approach outperforms fine-tuning and prior steering methods on both LLM and VLM tasks without degrading fluency or general capabilities.
A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for fine-tuning large language models to meet safety policies are costly and impractical.
Approach: They propose a method to fine-tune large language models to meet evolving safety policies by applying a gradient-free, unsupervised approach.
Outcome: The proposed method provides precise control, avoids blanket refusals, and directs models to generate safe, relevant content.
Multi-Attribute Steering of Language Models via Targeted Intervention (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches for steering large language models fail to scale to multi-attribute settings with conflicts, such as enhancing helpfulness while also reducing toxicity.
Approach: They propose a steering framework for selective token-level intervention across multiple attributes that enforcing sparsity and orthogonality among vectors for different attributes.
Outcome: The proposed framework outperforms existing ITI and parameter-efficient fine-tuning approaches across question answering tasks and generative tasks.
Representation Bending for Large Language Model Safety (2025.acl-long)

Copied to clipboard

Challenge: Existing safety-enhancing techniques, such as fine-tuning with human feedback or adversarial training, are still vulnerable as they address specific threats and fail to generalize across unseen attacks.
Approach: They propose a new approach that disrupts representations underlying harmful behaviors in Large Language Models by using loss-based fine-tuning.
Outcome: The proposed approach outperforms existing methods such as Circuit Breaker, RMU, and NPO with 95% reduction in attack success rates across diverse jailbreak benchmarks.
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering (2025.findings-acl)

Copied to clipboard

Challenge: Existing prompt-based debiasing methods exhibit instability due to sensitivity to prompt changes . fine-tuning-based techniques incur substantial computational overhead and catastrophic forgetting .
Approach: They propose a debiasing framework that encodes fairness-related features into separable directions in the hidden activation space.
Outcome: The proposed framework performs inference-time debiasing without requiring retraining or prompt design . it detects bias signatures in activations and then computes debiased steering vectors . the proposed framework is available to download in the u.s.
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Despite efforts to mitigate social bias in large language models, representational harms such as stereotyping continue to exist in both open and closed-source models.
Approach: They propose a method to modify model activations in forward passes by applying steering vectors to a BBQ dataset and comparing their results to bias mitigation methods.
Outcome: The proposed method outperforms 3 other bias mitigation methods on the BBQ dataset and shows the lowest impact on MMLU scores.
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for inference-time steering are limited by their limitations . Angular Steering violates norm preservation, causing distribution shift and generation collapse .
Approach: They propose a method that uses a norm-preserving rotation formulation to maintain activation distribution integrity and discriminative layer selection to apply steering only where features exhibit opposite-signed class alignment.
Outcome: Experiments show that Selective Steering achieves higher attack success rates than prior methods while maintaining zero perplexity violations and approximately 100% capability retention on standard benchmarks.
Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms (2025.acl-long)

Copied to clipboard

Challenge: Recent research has explored the use of sparse autoencoders (SAE) to disentangle knowledge in high-dimensional spaces for steering.
Approach: They propose a method that isolates and manipulates disentangled knowledge components to enhance safety by using sparse autoencoders to disentangle knowledge in high-dimensional spaces for steering.
Outcome: The proposed method is able to isolate and manipulate disentangled knowledge components to enhance safety in large reasoning models.
SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities in natural language understanding and generation, but controlling their behavior remains a challenge.
Approach: They propose a supervised steering approach that operates in sparse, interpretable representation spaces.
Outcome: The proposed approach achieves higher success rates with minimal degradation in generation quality compared to existing methods.
Can Activation Steering Generalize Across Languages? A Study on Syllogistic Reasoning in Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Prior work has focused on activation steering for Large Language Models (LLMs) this technique can be used to improve reasoning accuracy and transferability across languages.
Approach: They propose to use activation steering to steer models towards a cross-lingual reasoning space.
Outcome: The proposed techniques generalise well to multilingual datasets while minimizing language modelling performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations