Papers by Shipeng Wang
SAME: Safety-Aware Model Editing Guided by Safety Transformation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing models that update or insert new knowledge require sequential parameter updates while maintaining model capability. |
| Approach: | They propose a model editing approach that estimates safety transforms and identifies corresponding safety direction in the neural activation space and aligns neural activations and network parameter updates under the safety constraints. |
| Outcome: | The proposed approach reduces unsafe responses to malicious queries while preserving the effectiveness of model editing. |