Papers with CBRN
Nuclear Deployed!: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are evolving into autonomous decision-makers, raising concerns about catastrophic risks in high-stakes domains, particularly in Chemical, Biological, Radiological and Nuclear (CBRN) . |
| Approach: | They propose a framework that is carefully constructed to effectively and naturally expose catastrophic risks in high-stakes domains such as CBRN. |
| Outcome: | The proposed framework exposes LLM agents to catastrophic behaviors and deception without being deliberately induced. |
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility (2025.emnlp-main)
Copied to clipboard
Brendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng, Julius Broomfield, Adam Gleave, Kellin Pelrine
| Challenge: | a recent study shows that fine-tuning can produce helpful-only models with safeguards destroyed. |
| Approach: | They propose a method for fine-tuning models to generate detailed, high-quality responses to harmful requests. |
| Outcome: | The proposed method produces helpful-only models with safeguards destroyed . OpenAI, Google, and Anthropic models will fully comply with requests for CBRN assistance . |