DeVisE: Towards the Behavioral Testing of Medical Large Language Models (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing evaluations of large language models do not reveal whether their outputs reflect genuine medical reasoning or superficial correlations. |
| Approach: | They propose a framework that probes fine-grained clinical understanding through controlled counterfactuals. |
| Outcome: | The proposed framework is based on demographic and vital signs data from the ICU discharge notes of patients in the intensive care unit (MIMIC-IV). |
Similar Papers
To Err Is Human, How about Medical Large Language Models? Comparing Pre-trained Language Models for Medical Assessment Errors and Reliability (2024.lrec-main)
Copied to clipboard
| Challenge: | a 1999 report found that at least forty thousand deaths are a result of preventable medical errors. |
| Approach: | They test pre-trained language models to characterize their error generation and reliability in medical assessment ability. |
| Outcome: | The results show that pre-trained models can generate errors and perform better than human models. |
How Can We Diagnose and Treat Bias in Large Language Models for Clinical Decision-Making? (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent studies have shown that LLMs exhibit social biases inherited from training data. |
| Approach: | They propose a framework for evaluation and mitigation of bias in Large Language Models applied to complex clinical cases using a dataset based on the JAMA Clinical Challenge. |
| Outcome: | The proposed framework employs multiple choice questions and explanations to evaluate gender and ethnicity biases in LLMs. |
How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations of Large Language Models for clinical numerical reasoning provide limited operation-level coverage and limited robustness of numerical understanding across clinical note formats. |
| Approach: | They propose a benchmarking tool that evaluates four main types of clinical numeracy . they present longitudinal MIMIC-IV vital-sign records in three semantically equivalent representations . |
| Outcome: | The proposed benchmark evaluates four main types of clinical numeracy: value retrieval, arithmetic computation, relational comparison, and aggregation. |
Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective (2025.coling-main)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) are not effective in solving real-world healthcare tasks, but they are able to provide demographic information and provide biased health predictions. |
| Approach: | They evaluate state-of-the-art LLMs with three prevalent learning frameworks across six diverse healthcare tasks and find significant challenges in applying LLM to real-world healthcare tasks. |
| Outcome: | The proposed models perform poorly in real-world healthcare tasks and are inconsistent with existing learning frameworks. |
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings. |
| Approach: | They propose safety guidelines for the potential deployment of large language models for mental health response. |
| Outcome: | The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory. |
Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Rapid medical concept drift can lead LLMs to provide incorrect or outdated advice. |
| Approach: | They propose to evaluate how large language models manage knowledge conflicts in clinical guidelines. |
| Outcome: | The proposed benchmark evaluates how LLMs manage varied knowledge conflicts in clinical guidelines. |
MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations test factual medical knowledge in isolation or assess patient-level reasoning without verifying correctness, leaving a critical gap. |
| Approach: | They propose a benchmark that links MIMIC-IV EHRs to a unified knowledge base built from UMLS and other biomedical vocabularies. |
| Outcome: | The proposed model improves by +16.4 macro-F1 points over the base model and eliminates truth inversion errors. |
Can LLMs Reason Like Doctors? Exploring the Limits of Large Language Models in Complex Medical Reasoning (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown remarkable progress in reasoning across multiple domains, but it remains unclear whether their abilities reflect genuine reasoning or sophisticated pattern matching. |
| Approach: | They conduct one of the largest evaluations to date, assessing 77 LLMs . they select three medical question answering (QA) benchmarks targeting reasoning processes . |
| Outcome: | The results highlight the need to improve specific reasoning strategies to better reflect medical decision-making. |
Med-CoDE: Medical Critique based Disagreement Evaluation Framework (2025.naacl-srw)
Copied to clipboard
| Challenge: | Existing evaluation methods for large language models lack robustness and accuracy in medical contexts. |
| Approach: | They propose an evaluation framework for medical LLMs that measures disagreement between model-generated responses and established medical ground truths. |
| Outcome: | The proposed evaluation framework captures accuracy and reliability in medical settings. |
Can LLMs Replace Clinical Doctors? Exploring Bias in Disease Diagnosis by Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a new study examines the bias of disease prediction in large language models . the model biases are prevalent across gender, age range and disease judgment behaviors . |
| Approach: | They propose a prompt-based approach to alleviate the bias in disease prediction with LLMs. |
| Outcome: | The proposed model alleviates the observed bias in disease prediction with LLMs. |