Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are proving significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. |
| Approach: | They propose a framework that deconstructs benchmark development into five stages from design to governance and provides a checklist of 46 medically-tailored criteria. |
| Outcome: | The framework deconstructs benchmark development into five stages from design to governance and provides a comprehensive checklist of 46 medically-tailored criteria. |
Similar Papers
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation (2025.findings-emnlp)
Copied to clipboard
Ming Zhang, Yujiong Shen, Zelin Li, Huayu Sha, Binze Hu, Yuhui Wang, Chenhao Huang, Shichun Liu, Jingqi Tong, Changhao Jiang, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang
| Challenge: | Current medical benchmarks have limitations in question design, data sources and evaluation methods. |
| Approach: | They propose a new benchmark covering five core medical areas . it includes 2,996 questions created from real-world electronic health records . |
| Outcome: | The proposed model covers five core medical areas and includes 2,996 questions created from real-world electronic health records and expert-designed clinical scenarios. |
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations (2025.emnlp-main)
Copied to clipboard
Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, null Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao
| Challenge: | Existing benchmarks assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments. |
| Approach: | They propose a new step-by-step evaluation pipeline that assesses formula selection, entity extraction, and arithmetic computation. |
| Outcome: | The proposed method improves the accuracy of large language models on medical benchmarks from 16.35% to 53.19%. |
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings (2026.eacl-industry)
Copied to clipboard
Jean-Philippe Corbeil, Minseon Kim, Maxime Griot, Sheela Agarwal, Alessandro Sordoni, Francois Beaulieu, Paul Vozila
| Challenge: | Existing risk evaluations focused on general safety benchmarks, resulting in role-dependent vulnerabilities in real-world medical and clinical deployments. |
| Approach: | They propose a patient-oriented dataset called PatientSafetyBench that evaluates a variety of open- and closed-source LLMs. |
| Outcome: | The proposed benchmark examines medical risks from 466 open- and closed-source LLMs across 5 risk categories. |
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes (2025.findings-acl)
Copied to clipboard
| Challenge: | Several studies have shown that large language models can answer medical questions correctly, outperforming the average human score in some medical exams. |
| Approach: | They introduce MEDEC, the first publicly available benchmark for medical error detection and correction in clinical notes. |
| Outcome: | The proposed model outperforms medical doctors in errors detection and correction tasks. |
MedEthicEval: Evaluating Large Language Models Based on Chinese Medical Ethics (2025.naacl-industry)
Copied to clipboard
| Challenge: | Large language models (LLMs) have been used in clinical decision support, medical education and patient communication. |
| Approach: | They propose a benchmark to evaluate large language models in the domain of medical ethics and assess their grasp of medical ethical principles and their application across diverse scenarios. |
| Outcome: | The proposed framework assesses the models’ grasp of medical ethics principles and their ability to apply them across diverse scenarios. |
Benchmarking LLMs on Authentic Cases from Medical Journals (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing medical benchmarks suffer from performance saturation due to medical exam questions. |
| Approach: | They evaluate the performance of over 20 open-source and proprietary large language models and benchmark them against human medical experts. |
| Outcome: | The new benchmark is based on authentic clinical cases sourced from medical journals and implements rigorous human review process to ensure the quality and reliability of the benchmark. |
CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios (2024.emnlp-main)
Copied to clipboard
| Challenge: | Chinese medical large language models (LLMs) are underperforming on this benchmark, especially where medical reasoning and factual consistency are vital. |
| Approach: | They propose a benchmark with 14 expert-guided clinical scenarios to assess the medical ability of large language models across 7 pivot dimensions. |
| Outcome: | The proposed benchmark has been validated in several ways. |
EMPEC: A Comprehensive Benchmark for Evaluating Large Language Models Across Diverse Healthcare Professions (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) show their potential in accurately answering biomedical questions, yet current healthcare benchmarks primarily assess knowledge mastered by medical doctors, neglecting other essential professions. |
| Approach: | They evaluated 17 LLMs including proprietary and open-source models and found they struggled with specialized fields and alternative medicine. |
| Outcome: | The examinations for medical PErsonnel in Chinese (EMPEC) features 157,803 exam questions across 124 subjects and 20 healthcare professions. |
MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations (2026.findings-acl)
Copied to clipboard
Congbo Ma, Yichun Zhang, Yousef Al-Jazzazi, Ahamed Foisal, Laasya Sharma, Yousra Sadqi, Khaled Saleh, Jihad Mallat, Farah E. Shamout
| Challenge: | Existing or generated clinical text may contain inaccuracies that can lead to serious adverse outcomes. |
| Approach: | They introduce a multilingual benchmark for error detection, localization and correction . they assessed the performance of a range of general-purpose, language-specific, and medical-domain language models . |
| Outcome: | The proposed benchmark covers English, Arabic and Chinese, with natural medical cases annotated and reviewed by domain experts. |
MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model Evaluation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing medical datasets require high quality domain-specific datasets. |
| Approach: | They propose a multi-level, multi-task, and multi-domain medical benchmark to facilitate the development of language models for healthcare. |
| Outcome: | The proposed model provides granular potential usage and supports a wide range of tasks. |