Papers by Nirav Diwan
Fingerprinting Fine-tuned Language Models in the Wild (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing fingerprinting methods to fingerprint language models are limited to attributing organic text . however, fine-tuned LMs can generate long, coherent, and grammatically valid synthetic text. |
| Approach: | They conduct extensive experiments to demonstrate the limitations of existing fingerprinting approaches. |
| Outcome: | The proposed fingerprinting methods are limited to attributing synthetic text generated by 10 pre-trained LMs. |
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? (2025.findings-emnlp)
Copied to clipboard
Muntasir Wahed, Xiaona Zhou, Kiet A. Nguyen, Tianjiao Yu, Nirav Diwan, Gang Wang, Dilek Hakkani-Tür, Ismini Lourentzou
| Challenge: | Recent advances in Large Language Models have significantly enhanced their code generation capabilities, but their robustness against adversarial misuse remains underexplored. |
| Approach: | They introduce a code decomposition attack where a malicious coding task is broken down into subtasks across multiple conversational turns to evade safety filters. |
| Outcome: | The proposed code decomposition attacks exploits multi-turn malicious coding prompts . the proposed model improves rejection rates while preserving coding ability . |