Papers by Govind Ramesh
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using large language models to jailbreak is important for testing safety and security issues. |
| Approach: | They propose an approach that leverages the reflective capabilities of large language models for jailbreaking with only black-box access. |
| Outcome: | The proposed method achieves jailbreak success rates of 98% on GPT-4, 92% on GTP-4 Turbo, and 94% on Llama-3.1-70B in under 7 queries. |