Papers by Jiamu Zhang
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Assistant Scenarios (2025.findings-acl)
Copied to clipboard
Jun Wang, Jiamu Zhou, Xihuai Wang, Xiaoyun Mo, Haoyu Zhang, Qiqiang Lin, Jincheng Jincheng, Muning Wen, Weinan Zhang, Qiuying Peng, Jun Wang
| Challenge: | Evaluating the performance of LLMs in multi-turn interactions presents significant challenges due to the complexity and variability of user behavior. |
| Approach: | They propose a benchmark framework for assessing LLMs’ function-calling capabilities in multi-turn dialogues. |
| Outcome: | The proposed framework is based on a dataset derived from popular mobile apps and anonymized user logs. |
ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution (2026.acl-industry)
Copied to clipboard
Jihong Wang, Jiamu Zhou, Weiming Zhang, Teng Wang, Weiwen Liu, Zhuosheng Zhang, Xingyu Lou, Weinan Zhang, Huarong Deng, Jun Wang
| Challenge: | Xue et al., 2025): deploying autonomous web agents in production remains difficult due to site heterogeneity and long-horizon instability. |
| Approach: | They propose a knowledge-evolving agent that can be used to automate web workflows . they use human-in-the-loop knowledge adaptation and knowledge-aligned progressive summarization . |
| Outcome: | Experiments on WebArena, WebChoreAren and industrial deployment show it outperforms baselines. |
ReasonerRank: Redefining Language Model Evaluation with Ground-Truth-Free Ranking Frameworks (2025.findings-acl)
Copied to clipboard
Jiamu Zhang, Jiayi Yuan, Andrew Wen, Hoang Anh Duy Le, Yu-Neng Chuang, Soo-Hyun Choi, Rui Chen, Xia Hu
| Challenge: | Large Language Models (LLMs) are increasingly adopted across real-world applications . traditional evaluations rely on expensive, domain-specific ground-truth labels . obtaining labeled data is expensive, time-consuming, and often requires domain expertise . |
| Approach: | They propose a ground-truth-free evaluation framework focused on reasoning consistency and instruction following. |
| Outcome: | The proposed framework outperforms existing label-free methods, including majority voting, triplet ranking, and peer-review approaches. |