Papers by Lingfeng Zhong
Activation Decomposition and Steering for LLM Backdoor Remediation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to defending against LLM backdoors rely on auxiliary models or safety-related datasets. |
| Approach: | They propose a method which contrasts benign and poisoned settings to decompose feature vectors for steering without auxiliary models or datasets. |
| Outcome: | The proposed method achieves better defense qualities than existing steering strategies. |