Papers by Yakun Zhu
MMXU: A Multi-Modal and Multi-X-ray Understanding Dataset for Disease Progression (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets and models fail to consider critical aspects of medical diagnostics, authors argue . MMXU enables multi-image questions incorporating both current and historical patient data. |
| Approach: | They propose a dataset for MedVQA that focuses on identifying changes in specific regions between two patient visits. |
| Outcome: | The proposed dataset improves diagnostic accuracy by 20% by integrating historical data. |
Perception, Understanding and Reasoning: A Multimodal Benchmark for Video Fake News Detection (2026.acl-long)
Copied to clipboard
| Challenge: | Existing video fake news detection benchmarks focus on the detection accuracy, while failing to provide fine-grained assessments for the entire detection process. |
| Approach: | They propose a process-oriented video fake news detection benchmark that evaluates MLLMs' perception, understanding, and reasoning capabilities in VFND. |
| Outcome: | The proposed model achieves sota performance on video fake news detection tasks. |
InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents (2026.findings-acl)
Copied to clipboard
Zhenghao Zhu, Yuanfeng Song, Xing Chen, Chengzhong Liu, Cui Yakun, Caleb Chen Cao, Sirui Han, Yike Guo
| Challenge: | Existing frameworks for data analysis and insight exploration are lacking in terms of benchmarks . existing frameworks suffer from format inconsistencies, poorly conceived objectives, and redundant insights. |
| Approach: | They propose a data-curation pipeline to construct a new dataset named InsightEval. |
| Outcome: | The proposed benchmarks highlight prevailing challenges in automated insight discovery and raise key findings to guide future research. |
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models (2026.findings-acl)
Copied to clipboard
Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
| Challenge: | Existing medical benchmarks for diagnostic reasoning are limited in their ability to perform complex tasks. |
| Approach: | They propose to benchmark diagnostic capabilities of large language models to assess their accuracy and generalization bottlenecks. |
| Outcome: | The proposed model achieves 45.82%, 31.09%, and 17.79% accuracy, compared to current models, o3-mini, e1 and DeepSeek-R1 . |
Meta-Tool: Unleash Open-World Function Calling Capabilities of General-Purpose Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models struggle with addressing diverse user inquiries in open-world tasks. |
| Approach: | They propose a plug-and-play tool retrieval system for LLMs to access external tool library and use retrieved tools to solve user's problem. |
| Outcome: | The proposed model improves on a finetuned version of LLaMA-3.1 and 2,800 dialogues and 7,361 tools spanning ten distinct test categories. |
MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling (2025.naacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been widely used in medicine but are limited in their ability to fully address the complexities of the real world. |
| Approach: | They propose a universal agent architecture for Large Language Models that integrates a specialized medical toolkit and employs meta-tool and nested calling mechanisms to enhance LLM tool utilization. |
| Outcome: | The proposed framework improves the accuracy and performance of medical calculators in complex medical scenarios. |
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on static single-step calculations with explicit instructions. |
| Approach: | They propose a benchmark for evaluating medical calculators in realistic scenarios . they use 118 scenario tasks across 4 clinical domains to evaluate medical calculator performance . |
| Outcome: | The first benchmark for evaluating medical calculators in realistic scenarios is released . it features 118 scenario tasks across 4 clinical domains and is based on a model context protocol integration. |