Papers by Jason Chu
LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion (2025.acl-long)
Copied to clipboard
| Challenge: | Existing safety alignment methods rely on fine-tuning, which inadvertently leads to the increased complexity and computational resources required. |
| Approach: | They propose a safety re-alignment framework with Low-Rank Safety Subspace Fusison that exploits low-rank safety characteristics of LLMs by constructing a low-ranked projection matrix to extract the principal components of safety vectors. |
| Outcome: | The proposed method exploits low-rank safety subspace of the LLMs and is stable during fine-tuning process and is isolated from the model’s general capabilities. |