Papers by Eitan Farchi

3 papers
Evaluating the Prompt Steerability of Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a primary question underlying alignment research is: whose views are we aligning to?
Approach: They propose to evaluate the steerability of model personas as a function of prompting by defining a benchmark and inspecting how these indices change as if steering effort is a factor.
Outcome: The proposed benchmark reveals that the steerability of many current models is limited due to skew in baseline behavior and an asymmetry in their steerability across many persona dimensions.
Exploring Straightforward Methods for Automatic Conversational Red-Teaming (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used in business dialogue systems but they also pose security and ethical risks.
Approach: They propose to use off-the-shelf large language models to create red-team attacks by eliciting undesired outputs from an attacker LLM.
Outcome: The proposed models can adapt their attack strategies based on prior attempts, but their effectiveness decreases as the alignment of the target model improves.
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios (2024.findings-emnlp)

Copied to clipboard

Challenge: Using large language models, we evaluated their robustness on multiple datasets.
Approach: They propose a new metric for assessing model robustness by empirical evaluation of several models on multiple datasets.
Outcome: The proposed metric is based on a set of datasets that are constructed by introducing naturally-occurring, non-malicious perturbations or by generating semantically equivalent paraphrases of input questions or statements.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations