Papers by Zhaochen Hong
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities (2025.acl-long)
Copied to clipboard
| Challenge: | Traditional self-consistency methods fail to capture subtle semantic errors in multi-step tasks. |
| Approach: | They propose a tree-based evaluation framework that measures LLMs’ ability to preserve semantic consistency during reversible transformations. |
| Outcome: | The proposed framework measures generalization abilities across models from 1.5B to 72B and can be used to benchmark LLMs without constructing new datasets. |
MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents (2025.acl-long)
Copied to clipboard
Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Robert Tang, Heng Ji, Jiaxuan You
| Challenge: | Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. |
| Approach: | They propose a benchmark to evaluate LLM-based multi-agent systems across diverse, interactive scenarios. |
| Outcome: | The proposed framework measures task completion and quality of collaboration and competition using novel, milestone-based key performance indicators. |