Papers by Roy Bar-Haim
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Evaluating debate speeches requires a deep understanding of arguments at multiple levels. |
| Approach: | They propose a benchmark task for LLM judges based on annotated debate speeches . they analyze the judgment capabilities and behavior of frontier LLMs . |
| Outcome: | The proposed task requires a comprehensive understanding of argumentation and its arguments. |
From Surrogacy to Adoption; From Bitcoin to Cryptocurrency: Debate Topic Expansion (P19-1)
Copied to clipboard
Roy Bar-Haim, Dalia Krieger, Orith Toledo-Ronen, Lilach Edelstein, Yonatan Bilu, Alon Halfon, Yoav Katz, Amir Menczel, Ranit Aharonov, Noam Slonim
| Challenge: | Recent advances in argumentation mining have left much of the relevant argumentative content out of reach. |
| Approach: | They propose a task of Debate Topic Expansion to find related topics for a given debate topic, along with an annotated dataset for the task. |
| Outcome: | The proposed algorithms differ from well-studied lexical-semantic relations and show they work well in argumentation mining. |
From Arguments to Key Points: Towards Automatic Argument Summarization (2020.acl-main)
Copied to clipboard
| Challenge: | Recent work on topic-related argument mining has made it difficult to read and digest large amounts of information. |
| Approach: | They propose to represent arguments as a small set of talking points, termed key points, each scored according to its salience. |
| Outcome: | The proposed method can predict key points in advance, and it performs well. |
Welcome to the Real World: Efficient, Incremental and Scalable Key Point Analysis (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Key Point Analysis (KPA) extracts the main points from opinions and quantifies their prevalence. |
| Approach: | They propose a key point analysis framework that extracts the main points from opinions and quantifies their prevalence. |
| Outcome: | The proposed system is able to match sentences to key points over five datasets and demonstrate its performance. |
Quantitative argument summarization and beyond: Cross-domain key point analysis (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent work on multi-document summarization lacks quantitative aspect of summarizing views, arguments or opinions . authors develop method for automatic extraction of key points, which is comparable to a human expert . |
| Approach: | They propose to map arguments to a small set of expert-generated key points . they demonstrate that the applicability of key point analysis goes well beyond argumentation data . |
| Outcome: | The proposed method outperforms arguments in municipal surveys and user reviews . it is shown that the extraction of key points is comparable to a human expert . |
Advances in Debating Technologies: Building AI That Can Debate Humans (2021.acl-tutorials)
Copied to clipboard
| Challenge: | This tutorial focuses on Debating Technologies, a sub-field of computational argumentation defined as "computational technologies developed directly to enhance, support, and engage with human debating" the tutorial provides a holistic view of a debated system, and discusses practical applications and future challenges of debation technologies. |
| Approach: | They present a tutorial on Debating Technologies, a sub-field of computational argumentation . they introduce Project Debater, which is the first AI system to debate human experts . |
| Outcome: | The project Debater is the first AI system to debate human experts on complex topics. |
JuStRank: Benchmarking LLM Judges for System Ranking (2025.acl-long)
Copied to clipboard
| Challenge: | Recent work has focused on instance-based evaluation of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems. |
| Approach: | They propose to validate the quality of the LLM judge itself by comparing system scores to a human-based ranking. |
| Outcome: | The proposed model fails to validate the quality of the judge itself, ignoring critical factors affecting system-level ranking, such as a judge’s positive or negative bias towards certain systems. |
From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization (2023.acl-long)
Copied to clipboard
| Challenge: | Key Point Analysis (KPA) is a new method for analyzing textual comments . it uses a list of concise sentences or phrases to extract key points from data . |
| Approach: | They propose to organize key points into a hierarchy according to their specificity . they compare methods for predicting pairwise relations between key points . |
| Outcome: | The proposed method improves on predicting pairwise key point relations and weak supervision. |
A Survey on Evaluation of LLM-based Agents (2026.findings-acl)
Copied to clipboard
Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, Michal Shmueli-Scheuer
| Challenge: | This paper provides the first comprehensive survey of evaluation methods for LLM-based agents . LLMs are static, having fixed knowledge, and confined to text-to-text interaction. |
| Approach: | They analyze the evaluation of LLM-based agents across five perspectives . they identify current trends and key gaps in evaluation methods . |
| Outcome: | The proposed evaluation frameworks and tools are based on five perspectives . the results highlight current trends and identify gaps in future research . |
Learning Sentiment Composition from Sentiment Lexicons (C18-1)
Copied to clipboard
Orith Toledo-Ronen, Roy Bar-Haim, Alon Halfon, Charles Jochim, Amir Menczel, Ranit Aharonov, Noam Slonim
| Challenge: | Sentiment composition is a fundamental problem in sentiment analysis. |
| Approach: | They propose a method for learning sentiment composition from a large, unlabeled corpus . they automatically generate large sentiment lexicons of bigrams and unigrams . |
| Outcome: | The proposed approach is validated through manual annotation and sentiment classification experiments with phrase-level and sentence-level benchmarks. |
Project Debater APIs: Decomposing the AI Grand Challenge (2021.emnlp-demo)
Copied to clipboard
| Challenge: | Project Debater is the first AI system that can debate human experts on complex topics. |
| Approach: | They describe Project Debater's architecture and evaluate its performance . they will focus on Key Point Analysis, a novel technology that identifies main points . |
| Outcome: | The proposed system can debate human experts on complex topics. |
Every Bite Is an Experience: Key Point Analysis of Business Reviews (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods for summarizing reviews focus on measuring sentiment toward aspects of the review . recent work shows that KPA improves performance without domain-specific annotation . |
| Approach: | They propose a framework that provides both textual and quantitative summary of the main points in review data. |
| Outcome: | The proposed framework significantly improves on existing methods without annotations and human supervision. |
CHAMP: Efficient Annotation and Consolidation of Cluster Hierarchies (2023.emnlp-demo)
Copied to clipboard
| Challenge: | Various annotation tasks require a complex hierarchical structure over nodes, where each node is a cluster of items. |
| Approach: | They propose an open source tool that incrementally constructs clusters and hierarchy simultaneously over any type of text. |
| Outcome: | The proposed approach significantly reduces annotation time and guarantees transitivity at the cluster and hierarchy levels. |
SLIDE - a Sentiment Lexicon of Common Idioms (L18-1)
Copied to clipboard
| Challenge: | Compositional solutions for phrase sentiment are not able to handle idioms because their sentiment is not derived from the sentiment of the individual words. |
| Approach: | They propose a crowdsourcing approach for collecting sentiment annotations of idiomatic expressions using crowdsourcing. |
| Outcome: | The proposed approach is able to capture sentiment strength and ambiguity in idiomatic expressions using crowdsourcing. |