Papers by Brendan Murphy
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility (2025.emnlp-main)
Copied to clipboard
Brendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng, Julius Broomfield, Adam Gleave, Kellin Pelrine
| Challenge: | a recent study shows that fine-tuning can produce helpful-only models with safeguards destroyed. |
| Approach: | They propose a method for fine-tuning models to generate detailed, high-quality responses to harmful requests. |
| Outcome: | The proposed method produces helpful-only models with safeguards destroyed . OpenAI, Google, and Anthropic models will fully comply with requests for CBRN assistance . |