Papers by Tom Tseng

1 papers
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study shows that fine-tuning can produce helpful-only models with safeguards destroyed.
Approach: They propose a method for fine-tuning models to generate detailed, high-quality responses to harmful requests.
Outcome: The proposed method produces helpful-only models with safeguards destroyed . OpenAI, Google, and Anthropic models will fully comply with requests for CBRN assistance .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations