Challenge: Existing evaluations of large language models do not reveal whether their outputs reflect genuine medical reasoning or superficial correlations.
Approach: They propose a framework that probes fine-grained clinical understanding through controlled counterfactuals.
Outcome: The proposed framework is based on demographic and vital signs data from the ICU discharge notes of patients in the intensive care unit (MIMIC-IV).

Similar Papers

To Err Is Human, How about Medical Large Language Models? Comparing Pre-trained Language Models for Medical Assessment Errors and Reliability (2024.lrec-main)

Copied to clipboard

Challenge: a 1999 report found that at least forty thousand deaths are a result of preventable medical errors.
Approach: They test pre-trained language models to characterize their error generation and reliability in medical assessment ability.
Outcome: The results show that pre-trained models can generate errors and perform better than human models.
How Can We Diagnose and Treat Bias in Large Language Models for Clinical Decision-Making? (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have shown that LLMs exhibit social biases inherited from training data.
Approach: They propose a framework for evaluation and mitigation of bias in Large Language Models applied to complex clinical cases using a dataset based on the JAMA Clinical Challenge.
Outcome: The proposed framework employs multiple choice questions and explanations to evaluate gender and ethnicity biases in LLMs.
How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of Large Language Models for clinical numerical reasoning provide limited operation-level coverage and limited robustness of numerical understanding across clinical note formats.
Approach: They propose a benchmarking tool that evaluates four main types of clinical numeracy . they present longitudinal MIMIC-IV vital-sign records in three semantically equivalent representations .
Outcome: The proposed benchmark evaluates four main types of clinical numeracy: value retrieval, arithmetic computation, relational comparison, and aggregation.
Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective (2025.coling-main)

Copied to clipboard

Challenge: Existing large language models (LLMs) are not effective in solving real-world healthcare tasks, but they are able to provide demographic information and provide biased health predictions.
Approach: They evaluate state-of-the-art LLMs with three prevalent learning frameworks across six diverse healthcare tasks and find significant challenges in applying LLM to real-world healthcare tasks.
Outcome: The proposed models perform poorly in real-world healthcare tasks and are inconsistent with existing learning frameworks.
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings.
Approach: They propose safety guidelines for the potential deployment of large language models for mental health response.
Outcome: The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory.
Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Rapid medical concept drift can lead LLMs to provide incorrect or outdated advice.
Approach: They propose to evaluate how large language models manage knowledge conflicts in clinical guidelines.
Outcome: The proposed benchmark evaluates how LLMs manage varied knowledge conflicts in clinical guidelines.
MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations test factual medical knowledge in isolation or assess patient-level reasoning without verifying correctness, leaving a critical gap.
Approach: They propose a benchmark that links MIMIC-IV EHRs to a unified knowledge base built from UMLS and other biomedical vocabularies.
Outcome: The proposed model improves by +16.4 macro-F1 points over the base model and eliminates truth inversion errors.
Can LLMs Reason Like Doctors? Exploring the Limits of Large Language Models in Complex Medical Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) have shown remarkable progress in reasoning across multiple domains, but it remains unclear whether their abilities reflect genuine reasoning or sophisticated pattern matching.
Approach: They conduct one of the largest evaluations to date, assessing 77 LLMs . they select three medical question answering (QA) benchmarks targeting reasoning processes .
Outcome: The results highlight the need to improve specific reasoning strategies to better reflect medical decision-making.
Med-CoDE: Medical Critique based Disagreement Evaluation Framework (2025.naacl-srw)

Copied to clipboard

Challenge: Existing evaluation methods for large language models lack robustness and accuracy in medical contexts.
Approach: They propose an evaluation framework for medical LLMs that measures disagreement between model-generated responses and established medical ground truths.
Outcome: The proposed evaluation framework captures accuracy and reliability in medical settings.
Can LLMs Replace Clinical Doctors? Exploring Bias in Disease Diagnosis by Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: a new study examines the bias of disease prediction in large language models . the model biases are prevalent across gender, age range and disease judgment behaviors .
Approach: They propose a prompt-based approach to alleviate the bias in disease prediction with LLMs.
Outcome: The proposed model alleviates the observed bias in disease prediction with LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations