Papers by Ron Litman

2 papers
DREAM: Deep Research Evaluation with Agentic Metrics (2026.acl-long)

Copied to clipboard

Challenge: Recent benchmarks propose distinct methodologies, yet they suffer from the Mirage of Synthesis . static evaluators lack the tool-use capabilities required to assess temporal validity and factual correctness .
Approach: They propose a framework that instantiates the principle of capability parity by making evaluation agentic.
Outcome: The proposed framework is more sensitive to factual decay than existing benchmarks . large language models increasingly support autonomous, tool-using agents .
M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation (2024.naacl-short)

Copied to clipboard

Challenge: Document translation is a challenge for machine translation systems that focus on textual content at the sentence level, ignoring global context and visual layout structure.
Approach: They propose a benchmark dataset to evaluate document-level NMT systems . they use visual cues to preserve reading order and contiguous blocks of text .
Outcome: The proposed benchmarks assess document-level NMT systems on the comprehensive task of translating semi-structured documents.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations