Papers by Sindhu Padakandla

1 papers
SafeQuant: LLM Safety Analysis via Quantized Gradient Inspection (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to jailbreak Large Language Models (LLMs) use computationally intensive verification or require adversarial fine-tuning, leaving models vulnerable to advanced attacks.
Approach: They propose a framework that leverages quantized gradient patterns to identify harmful prompts efficiently.
Outcome: The proposed framework outperforms existing defenses across multiple benchmarks while maintaining model utility.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations