Papers by Andrew Bell
SAFENUDGE: Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are susceptible to jailbreak attacks, or adversarial prompts eliciting high-risk behavior. |
| Approach: | They propose a safeguard that combines Controlled Text Generation and "nudging" it adds minimal latency to inference and reduces successful jailbreak attempts by up to 37.3% . |
| Outcome: | The proposed safeguard reduces successful jailbreak attempts by between 28.1% and 37.3% by guiding the LLM towards a safe response. |