Papers by Eitan Farchi
Evaluating the Prompt Steerability of Large Language Models (2025.naacl-long)
Copied to clipboard
Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy, Elizabeth M. Daly, Kush R. Varshney, Eitan Farchi, Pierre Dognin, Jesus Rios, Djallel Bouneffouf, Miao Liu, Prasanna Sattigeri
| Challenge: | a primary question underlying alignment research is: whose views are we aligning to? |
| Approach: | They propose to evaluate the steerability of model personas as a function of prompting by defining a benchmark and inspecting how these indices change as if steering effort is a factor. |
| Outcome: | The proposed benchmark reveals that the steerability of many current models is limited due to skew in baseline behavior and an asymmetry in their steerability across many persona dimensions. |
Exploring Straightforward Methods for Automatic Conversational Red-Teaming (2025.naacl-industry)
Copied to clipboard
George Kour, Naama Zwerdling, Marcel Zalmanovici, Ateret Anaby Tavor, Ora Nova Fandina, Eitan Farchi
| Challenge: | Large language models (LLMs) are increasingly used in business dialogue systems but they also pose security and ethical risks. |
| Approach: | They propose to use off-the-shelf large language models to create red-team attacks by eliciting undesired outputs from an attacker LLM. |
| Outcome: | The proposed models can adapt their attack strategies based on prior attempts, but their effectiveness decreases as the alignment of the target model improves. |
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Using large language models, we evaluated their robustness on multiple datasets. |
| Approach: | They propose a new metric for assessing model robustness by empirical evaluation of several models on multiple datasets. |
| Outcome: | The proposed metric is based on a set of datasets that are constructed by introducing naturally-occurring, non-malicious perturbations or by generating semantically equivalent paraphrases of input questions or statements. |