Papers by Manying Zhang
On Weaponization-Resistant Large Language Models with Prospect Theoretic Alignment (2025.coling-main)
Copied to clipboard
| Challenge: | Existing safeguards for large language models are inadequate for open-weight models as minimal fine-tuning can bypass them. |
| Approach: | They propose a framework that prioritizes maximizing generative utility rather than a singular optimization metric and integrates prospect theory into LLM training to strengthen LLMs against misuse and weaponization. |
| Outcome: | The proposed framework strengthens LLMs against misuse and weaponization while maintaining high performance even after extensive fine-tuning. |