Challenge: Large language models (LLMs) have been used in clinical decision support, medical education and patient communication.
Approach: They propose a benchmark to evaluate large language models in the domain of medical ethics and assess their grasp of medical ethical principles and their application across diverse scenarios.
Outcome: The proposed framework assesses the models’ grasp of medical ethics principles and their ability to apply them across diverse scenarios.

Similar Papers

MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings (2026.eacl-industry)

Copied to clipboard

Challenge: Existing risk evaluations focused on general safety benchmarks, resulting in role-dependent vulnerabilities in real-world medical and clinical deployments.
Approach: They propose a patient-oriented dataset called PatientSafetyBench that evaluates a variety of open- and closed-source LLMs.
Outcome: The proposed benchmark examines medical risks from 466 open- and closed-source LLMs across 5 risk categories.
A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare (2025.findings-emnlp)

Copied to clipboard

Challenge: a survey of large language models in healthcare raises critical concerns around trustworthiness . trustworthy of LLMs in healthcare remains underexplored, lacking a systematic review .
Approach: a new survey examines the trustworthiness of large language models in healthcare . a review examines how each dimension affects reliability and ethical deployment of LLMs .
Outcome: The present study examines the trustworthiness of large language models in healthcare . it identifies key gaps in existing approaches and challenges posed by evolving paradigms .
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation (2025.findings-emnlp)

Copied to clipboard

Challenge: Current medical benchmarks have limitations in question design, data sources and evaluation methods.
Approach: They propose a new benchmark covering five core medical areas . it includes 2,996 questions created from real-world electronic health records .
Outcome: The proposed model covers five core medical areas and includes 2,996 questions created from real-world electronic health records and expert-designed clinical scenarios.
CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios (2024.emnlp-main)

Copied to clipboard

Challenge: Chinese medical large language models (LLMs) are underperforming on this benchmark, especially where medical reasoning and factual consistency are vital.
Approach: They propose a benchmark with 14 expert-guided clinical scenarios to assess the medical ability of large language models across 7 pivot dimensions.
Outcome: The proposed benchmark has been validated in several ways.
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are proving significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities.
Approach: They propose a framework that deconstructs benchmark development into five stages from design to governance and provides a checklist of 46 medically-tailored criteria.
Outcome: The framework deconstructs benchmark development into five stages from design to governance and provides a comprehensive checklist of 46 medically-tailored criteria.
Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on evaluating large language models in close-ended QA tasks, but many clinical decisions involve answering open-ended questions without pre-set options.
Approach: They construct a benchmark to better understand large language models in the clinic . they use existing datasets to evaluate LLMs in clinical situations .
Outcome: The proposed model outperforms human experts in multiple medical tasks.
Med-CoDE: Medical Critique based Disagreement Evaluation Framework (2025.naacl-srw)

Copied to clipboard

Challenge: Existing evaluation methods for large language models lack robustness and accuracy in medical contexts.
Approach: They propose an evaluation framework for medical LLMs that measures disagreement between model-generated responses and established medical ground truths.
Outcome: The proposed evaluation framework captures accuracy and reliability in medical settings.
Benchmarking LLMs on Authentic Cases from Medical Journals (2026.findings-acl)

Copied to clipboard

Challenge: Existing medical benchmarks suffer from performance saturation due to medical exam questions.
Approach: They evaluate the performance of over 20 open-source and proprietary large language models and benchmark them against human medical experts.
Outcome: The new benchmark is based on authentic clinical cases sourced from medical journals and implements rigorous human review process to ensure the quality and reliability of the benchmark.
Should I Believe in What Medical AI Says? A Chinese Benchmark for Medication Based on Knowledge and Reasoning (2025.acl-short)

Copied to clipboard

Challenge: Large language models (LLMs) generate hallucinations when handling unfamiliar information.
Approach: They propose a Chinese benchmark to evaluate large language models' knowledge and reasoning capabilities in medication tasks.
Outcome: The proposed benchmark evaluates models in indication, dosage and administration, contraindicated population, mechanisms of action, drug recommendation, and drug interaction across six datasets.
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments.
Approach: They propose a new step-by-step evaluation pipeline that assesses formula selection, entity extraction, and arithmetic computation.
Outcome: The proposed method improves the accuracy of large language models on medical benchmarks from 16.35% to 53.19%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations