Papers by Noy Sternlicht
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Evaluating debate speeches requires a deep understanding of arguments at multiple levels. |
| Approach: | They propose a benchmark task for LLM judges based on annotated debate speeches . they analyze the judgment capabilities and behavior of frontier LLMs . |
| Outcome: | The proposed task requires a comprehensive understanding of argumentation and its arguments. |
In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis (2026.acl-long)
Copied to clipboard
| Challenge: | citation counts are a shallow view that fails to capture how a paper has influenced subsequent work. |
| Approach: | They propose a task to generate nuanced, expressive, and time-aware impact summaries . they analyze fine-grained confirmatory and correction citation intents to generate summary . |
| Outcome: | The proposed task shows moderate to strong human correlation on subjective metrics such as insightfulness. |
CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation (2026.acl-long)
Copied to clipboard
| Challenge: | a hallmark of human innovation is recombination. |
| Approach: | They propose a task to extract recombination instances from scientific literature . they analyze patterns of recombined concepts and apply it to a broad corpus of AI papers . |
| Outcome: | The proposed model can predict cross-disciplinary research directions . it can predict recombinations across areas and link methods and concepts . |