Papers by Manying Zhang

1 papers
On Weaponization-Resistant Large Language Models with Prospect Theoretic Alignment (2025.coling-main)

Copied to clipboard

Challenge: Existing safeguards for large language models are inadequate for open-weight models as minimal fine-tuning can bypass them.
Approach: They propose a framework that prioritizes maximizing generative utility rather than a singular optimization metric and integrates prospect theory into LLM training to strengthen LLMs against misuse and weaponization.
Outcome: The proposed framework strengthens LLMs against misuse and weaponization while maintaining high performance even after extensive fine-tuning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations