Papers by Chongwen Zhao

1 papers
Unraveling LLM Jailbreaks Through Safety Knowledge Neurons (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved significant progress in alignment, ensuring safer and more reliable outputs.
Approach: They propose a neuron-level interpretability method that focuses on the role of safety-related knowledge neurons to improve model robustness against jailbreak attacks.
Outcome: The proposed method reduces attack success rates across multiple LLMs and outperforms all baseline defenses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations