Papers by Noah Flynn

2 papers
DREAM: Deep Research Evaluation with Agentic Metrics (2026.acl-long)

Copied to clipboard

Challenge: Recent benchmarks propose distinct methodologies, yet they suffer from the Mirage of Synthesis . static evaluators lack the tool-use capabilities required to assess temporal validity and factual correctness .
Approach: They propose a framework that instantiates the principle of capability parity by making evaluation agentic.
Outcome: The proposed framework is more sensitive to factual decay than existing benchmarks . large language models increasingly support autonomous, tool-using agents .
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Southeast Asia (SEA) is home to over 1,300 indigenous languages and 671 million people . prevailing AI models suffer from a significant lack of representation of texts, images, and audio datasets from SEA .
Approach: They propose to provide a resource center that provides standardized corpora in nearly 1,000 SEA languages across three modalities.
Outcome: a new benchmark assesses the quality of AI models on 36 SEA languages across 13 tasks . the results highlight the importance of SEA as a culturally diverse region .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations