Challenge: Existing methods to edit LLMs' activations are limited by their magnitude and direction consistency.
Approach: They propose a method that edits activations to alter their magnitudes and directions to preserve activation norms.
Outcome: The proposed method preserves activation norm and improves safety benchmarks.

Similar Papers

Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Prior work shows that Large Language Models exhibit highly anisotropic internal representations . prior work shows specialized dimensions capture domain-specific features .
Approach: They propose a simple magnitude-based criterion to identify Domain-Critical Dimensions in a training-free manner.
Outcome: The proposed method outperforms whole-dimension steering in domain adaptation and jailbreaking scenarios.
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Despite efforts to mitigate social bias in large language models, representational harms such as stereotyping continue to exist in both open and closed-source models.
Approach: They propose a method to modify model activations in forward passes by applying steering vectors to a BBQ dataset and comparing their results to bias mitigation methods.
Outcome: The proposed method outperforms 3 other bias mitigation methods on the BBQ dataset and shows the lowest impact on MMLU scores.
Activation Scaling for Steering and Interpreting Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: a successful intervention should flip the correct with the wrong token, while remaining sparse.
Approach: They propose to use activation scaling to flip the correct with the wrong token . they use gradient-based optimization to learn and evaluate a specific kind of efficient intervention .
Outcome: The proposed method performs comparable with steering vectors but is much less minimal.
Token-Aware Editing of Internal Activations for Large Language Model Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to optimize the behavior of large language models neglect misalignment discrepancies among tokens, resulting in deviant alignment direction and inflexible editing strength.
Approach: They propose a token-aware editing approach to exploit the misalignment discrepancy among tokens to enhance activation probing and facilitate intervention.
Outcome: Extensive experiments on three alignment capabilities demonstrate the efficacy of the proposed approach surpassing baseline by 25.8% on the primary metric of truthfulness with minimal cost.
Style Vectors for Steering Generative Large Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) can be trained on vast corpora and can generate text in a nuanced and parameterisable way.
Approach: They propose to add style vectors to the activations of hidden layers during text generation to steer output towards specific styles.
Outcome: The proposed approach differs from prompt engineering in that it can be nuanced and parameterisable.
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for inference-time steering are limited by their limitations . Angular Steering violates norm preservation, causing distribution shift and generation collapse .
Approach: They propose a method that uses a norm-preserving rotation formulation to maintain activation distribution integrity and discriminative layer selection to apply steering only where features exhibit opposite-signed class alignment.
Outcome: Experiments show that Selective Steering achieves higher attack success rates than prior methods while maintaining zero perplexity violations and approximately 100% capability retention on standard benchmarks.
Activation Steering Decoding: Mitigating Hallucination in Large Vision-Language Models through Bidirectional Hidden State Intervention (2025.acl-long)

Copied to clipboard

Challenge: Large Vision Language Models (LVLMs) suffer from hallucination where generated textual descriptions fail to align accurately with visual semantics.
Approach: They propose a training-free approach that mitigates hallucination through targeted intervention in the model’s intermediate activations by identifying directional patterns of hallucinism in the activation space using a small calibration set.
Outcome: The proposed approach reduces hallucination across multiple benchmarks while maintaining performance on general visual understanding tasks.
Internal Value Alignment in Large Language Models through Controlled Value Vector Activation (2025.acl-long)

Copied to clipboard

Challenge: Existing LLMs do not possess consistent values, but many have been developed to align them at the behavioral level, including supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF).
Approach: They propose a Controlled Value Vector Activation method that directly aligns the internal values of Large Language Models by interpreting how a value is encoded in their latent representations.
Outcome: The proposed method achieves highest success rate across 10 basic values without hurting model performance and fluency, and ensures target values even with opposite and potentially malicious input prompts.
Activation Decomposition and Steering for LLM Backdoor Remediation (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to defending against LLM backdoors rely on auxiliary models or safety-related datasets.
Approach: They propose a method which contrasts benign and poisoned settings to decompose feature vectors for steering without auxiliary models or datasets.
Outcome: The proposed method achieves better defense qualities than existing steering strategies.
Editing Large Language Models: Problems, Methods, and Opportunities (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in model editing for LLMs have created challenges and opportunities for the community.
Approach: They propose to alter the behavior of LLMs efficiently within a specific domain without negatively impacting performance across other inputs.
Outcome: The proposed method alters behavior of LLMs efficiently within a specific domain without negatively impacting performance across other inputs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations