Papers by Zhaochen Hong

2 papers
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities (2025.acl-long)

Copied to clipboard

Challenge: Traditional self-consistency methods fail to capture subtle semantic errors in multi-step tasks.
Approach: They propose a tree-based evaluation framework that measures LLMs’ ability to preserve semantic consistency during reversible transformations.
Outcome: The proposed framework measures generalization abilities across models from 1.5B to 72B and can be used to benchmark LLMs without constructing new datasets.
MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition.
Approach: They propose a benchmark to evaluate LLM-based multi-agent systems across diverse, interactive scenarios.
Outcome: The proposed framework measures task completion and quality of collaboration and competition using novel, milestone-based key performance indicators.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations