Papers by Alexandre Araujo
Stronger Universal and Transferable Attacks by Suppressing Refusals (2025.naacl-long)
Copied to clipboard
| Challenge: | Efforts have focused on aligning models to human preferences (RLHF) . yet, it is believed that such optimization-based attacks are sample-specific. |
| Approach: | They propose an algorithm to embed a "safety feature" into models to make them safe for mass deployment. |
| Outcome: | The proposed attack achieves 25% success rate against the state-of-the-art Circuit Breaker defense, compared to 2.5% by white-box GCG. |