Challenge: Systematic reviews should take into account the quality of available evidence, placing more weight on studies that use a valid methodology.
Approach: They propose to use a risk-of-bias framework to assess the methodological strength of biomedical papers by combining expert reviewers' judgments with research paper sentences.
Outcome: The proposed system measures the methodological strength of biomedical papers using the risk-of-bias framework used for systematic reviews.

Similar Papers

MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings (2026.eacl-industry)

Copied to clipboard

Challenge: Existing risk evaluations focused on general safety benchmarks, resulting in role-dependent vulnerabilities in real-world medical and clinical deployments.
Approach: They propose a patient-oriented dataset called PatientSafetyBench that evaluates a variety of open- and closed-source LLMs.
Outcome: The proposed benchmark examines medical risks from 466 open- and closed-source LLMs across 5 risk categories.
RoBGuard: Enhancing LLMs to Assess Risk of Bias in Clinical Trial Documents (2025.coling-main)

Copied to clipboard

Challenge: Existing approaches to assess the risk of bias in RCTs focus on manually crafted prompts and a restricted set of simple questions, limiting their accuracy and generalizability.
Approach: They propose a framework for enhancing Large Language Models to assess the risk of bias in RCTs by reformulation, document parsing and multi-expert collaboration.
Outcome: The proposed framework outperforms existing methods on the RoB-Item and RoB domains.
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments.
Approach: They propose a new step-by-step evaluation pipeline that assesses formula selection, entity extraction, and arithmetic computation.
Outcome: The proposed method improves the accuracy of large language models on medical benchmarks from 16.35% to 53.19%.
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are proving significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities.
Approach: They propose a framework that deconstructs benchmark development into five stages from design to governance and provides a checklist of 46 medically-tailored criteria.
Outcome: The framework deconstructs benchmark development into five stages from design to governance and provides a comprehensive checklist of 46 medically-tailored criteria.
Benchmarking LLMs on Authentic Cases from Medical Journals (2026.findings-acl)

Copied to clipboard

Challenge: Existing medical benchmarks suffer from performance saturation due to medical exam questions.
Approach: They evaluate the performance of over 20 open-source and proprietary large language models and benchmark them against human medical experts.
Outcome: The new benchmark is based on authentic clinical cases sourced from medical journals and implements rigorous human review process to ensure the quality and reliability of the benchmark.
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations (2023.acl-long)

Copied to clipboard

Challenge: Prior work has shown that models may exploit shortcuts that are difficult to detect using standard n-gram similarity metrics such as ROUGE.
Approach: They propose to use human-assessed summary quality facets and pairwise preferences to improve MDS evaluation methods.
Outcome: The proposed methods improve the quality of literature review summarization models . they use human-assessed summary quality facets and pairwise preferences .
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency (2025.emnlp-main)

Copied to clipboard

Challenge: Current dataset curation and bias assessment practices lack transparency . current approaches lack a thorough understanding of how data characteristics influence model behavior .
Approach: They propose a comprehensive bias evaluation framework that integrates general benchmarks with a healthcare-specific methodology to probe for biases in a sensitive healthcare context.
Outcome: The proposed approach to bias evaluation leverages established benchmarks and a healthcare-specific methodology.
Appraising the Potential Uses and Harms of LLMs for Medical Systematic Reviews (2023.emnlp-main)

Copied to clipboard

Challenge: Medical systematic reviews are time-consuming and often generate inaccurate outputs . authors: a model that generates scientific-sounding outputs can be unusable at best .
Approach: They conduct interviews with systematic review experts to characterize perceived utility and risks of LLMs in medical evidence reviews.
Outcome: a new study characterizes perceived utility and risks of medical evidence reviews . experts say they can assist in the writing process by drafting summaries, distilling information . authors say they expect the model to be more accurate and more reliable .
The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature (2022.acl-long)

Copied to clipboard

Challenge: Existing evaluation approaches to multi-document summarization of biomedical literature lack consistency and transparency.
Approach: They propose a systematic approach to human evaluation of biomedical summaries and apply it to analyze the summary generated by two current evaluation models.
Outcome: The proposed evaluation framework is based on two state-of-the-art models and examines the summaries generated by the two models to understand the deficiencies of existing evaluation approaches.
Curse of Knowledge: Your Guidance and Provided Knowledge are biasing LLM Judges in Complex Evaluation (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study has focused on simple settings, but their reliability in complex tasks remains understudied.
Approach: They propose to use large language models as judges to evaluate reliability in complex tasks . they use a challenge benchmark to expose and quantify Auxiliary Information Induced Biases .
Outcome: The proposed benchmark exposes and quantifies Auxiliary Information Induced Biases across 12 basic and 3 advanced scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations