Challenge: Clinical trials provide the highest quality evidence for clinical care . applying ROB2 is time-consuming, taking trained reviewers 30+ minutes per clinical trial report.
Approach: They propose an open-source platform for large language model-assisted risk of bias assessment of clinical trials.
Outcome: The proposed platform enables rapid and reproducible clinical trial annotations.

Similar Papers

RoBGuard: Enhancing LLMs to Assess Risk of Bias in Clinical Trial Documents (2025.coling-main)

Copied to clipboard

Challenge: Existing approaches to assess the risk of bias in RCTs focus on manually crafted prompts and a restricted set of simple questions, limiting their accuracy and generalizability.
Approach: They propose a framework for enhancing Large Language Models to assess the risk of bias in RCTs by reformulation, document parsing and multi-expert collaboration.
Outcome: The proposed framework outperforms existing methods on the RoB-Item and RoB domains.
Synthetic Data for Evaluation: Supporting LLM-as-a-Judge Workflows with EvalAssist (2025.emnlp-demos)

Copied to clipboard

Challenge: EvalAssist is a web-based application designed to assist human-centered evaluation of language model outputs.
Approach: They propose a synthetic data generation tool integrated into EvalAssist to assist human-centered evaluation of language model outputs.
Outcome: The proposed tool supports flexible prompting, RAG-based grounding, persona diversity, and iterative generation workflows.
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation (2026.eacl-demo)

Copied to clipboard

Challenge: Existing work on bias evaluation includes benchmark datasets and automated detection methods.
Approach: They propose an open-source web application for side-by-side comparison of LLM outputs with real-time bias analysis.
Outcome: The open-source application compares LLM outputs with real-time bias analysis.
LLM-Based Web Data Collection for Research Dataset Creation (2025.findings-emnlp)

Copied to clipboard

Challenge: researchers across many fields rely on web data to gain new insights and validate methods.
Approach: They propose a human-in-the-loop framework that automates web-scale data collection end-to-end using large language models (LLMs)
Outcome: The proposed framework outperforms existing methods in three different tasks and a user evaluation demonstrates its practical utility.
Humans or LLMs as the Judge? A Study on Judgement Bias (2024.emnlp-main)

Copied to clipboard

Challenge: Proprietary models such as GPT-4, Claude, Gemini-Pro and others are being democratized to improve evaluations of LLMs.
Approach: They propose a framework that is free from referencing groundtruth annotations for investigating **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia's** on LLM and human judges.
Outcome: The proposed framework investigates **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia' on LLM and human judges.
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments.
Approach: They propose a new step-by-step evaluation pipeline that assesses formula selection, entity extraction, and arithmetic computation.
Outcome: The proposed method improves the accuracy of large language models on medical benchmarks from 16.35% to 53.19%.
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency (2025.emnlp-main)

Copied to clipboard

Challenge: Current dataset curation and bias assessment practices lack transparency . current approaches lack a thorough understanding of how data characteristics influence model behavior .
Approach: They propose a comprehensive bias evaluation framework that integrates general benchmarks with a healthcare-specific methodology to probe for biases in a sensitive healthcare context.
Outcome: The proposed approach to bias evaluation leverages established benchmarks and a healthcare-specific methodology.
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown promising potential in the medical domain, assisting with tasks like clinical note generation and patient communication.
Approach: They propose a framework that excels at utilizing domain-specific tools within two stages.
Outcome: The proposed framework surpasses the pure LLMs with more than 10 points and the well-established agent-based methods with 3 points.
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation (2026.eacl-demo)

Copied to clipboard

Challenge: EvalSense is a flexible framework for constructing domain-specific evaluation suites for large language models . it provides out-of-the-box support for a broad range of model providers and evaluation strategies .
Approach: They propose a framework for constructing domain-specific evaluation suites for large language models.
Outcome: The proposed framework provides out-of-the-box support for a broad range of model providers and evaluation strategies.
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods based on large language models (LLMs) are expensive and lack expertise due to limitations in human expertise.
Approach: They propose an open-source automatic evaluation model with 13B parameters specifically engineered to measure the question-answering proficiency of medical LLMs.
Outcome: The proposed model surpasses baselines in terms of correlation with human judgments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations