Papers by Marcel Zalmanovici
Exploring Straightforward Methods for Automatic Conversational Red-Teaming (2025.naacl-industry)
Copied to clipboard
George Kour, Naama Zwerdling, Marcel Zalmanovici, Ateret Anaby Tavor, Ora Nova Fandina, Eitan Farchi
| Challenge: | Large language models (LLMs) are increasingly used in business dialogue systems but they also pose security and ethical risks. |
| Approach: | They propose to use off-the-shelf large language models to create red-team attacks by eliciting undesired outputs from an attacker LLM. |
| Outcome: | The proposed models can adapt their attack strategies based on prior attempts, but their effectiveness decreases as the alignment of the target model improves. |