ROBOTO2: An Interactive System and Dataset for LLM-assisted Clinical Trial Risk of Bias Assessment (2025.emnlp-demos)
Copied to clipboard
Anthony Hevia, Sanjana Chintalapati, Veronica Ka Wai Lai, Nguyen Thanh Tam, Wai-Tat Wong, Terry P Klassen, Lucy Lu Wang
| Challenge: | Clinical trials provide the highest quality evidence for clinical care . applying ROB2 is time-consuming, taking trained reviewers 30+ minutes per clinical trial report. |
| Approach: | They propose an open-source platform for large language model-assisted risk of bias assessment of clinical trials. |
| Outcome: | The proposed platform enables rapid and reproducible clinical trial annotations. |
Similar Papers
RoBGuard: Enhancing LLMs to Assess Risk of Bias in Clinical Trial Documents (2025.coling-main)
Copied to clipboard
Changkai Ji, Bowen Zhao, Zhuoyao Wang, Yingwen Wang, Yuejie Zhang, Ying Cheng, Rui Feng, Xiaobo Zhang
| Challenge: | Existing approaches to assess the risk of bias in RCTs focus on manually crafted prompts and a restricted set of simple questions, limiting their accuracy and generalizability. |
| Approach: | They propose a framework for enhancing Large Language Models to assess the risk of bias in RCTs by reformulation, document parsing and multi-expert collaboration. |
| Outcome: | The proposed framework outperforms existing methods on the RoB-Item and RoB domains. |
Synthetic Data for Evaluation: Supporting LLM-as-a-Judge Workflows with EvalAssist (2025.emnlp-demos)
Copied to clipboard
Martín Santillán Cooper, Zahra Ashktorab, Hyo Jin Do, Erik Miehling, Werner Geyer, Jasmina Gajcin, Elizabeth M. Daly, Qian Pan, Michael Desmond
| Challenge: | EvalAssist is a web-based application designed to assist human-centered evaluation of language model outputs. |
| Approach: | They propose a synthetic data generation tool integrated into EvalAssist to assist human-centered evaluation of language model outputs. |
| Outcome: | The proposed tool supports flexible prompting, RAG-based grounding, persona diversity, and iterative generation workflows. |
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation (2026.eacl-demo)
Copied to clipboard
| Challenge: | Existing work on bias evaluation includes benchmark datasets and automated detection methods. |
| Approach: | They propose an open-source web application for side-by-side comparison of LLM outputs with real-time bias analysis. |
| Outcome: | The open-source application compares LLM outputs with real-time bias analysis. |
LLM-Based Web Data Collection for Research Dataset Creation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | researchers across many fields rely on web data to gain new insights and validate methods. |
| Approach: | They propose a human-in-the-loop framework that automates web-scale data collection end-to-end using large language models (LLMs) |
| Outcome: | The proposed framework outperforms existing methods in three different tasks and a user evaluation demonstrates its practical utility. |
Humans or LLMs as the Judge? A Study on Judgement Bias (2024.emnlp-main)
Copied to clipboard
| Challenge: | Proprietary models such as GPT-4, Claude, Gemini-Pro and others are being democratized to improve evaluations of LLMs. |
| Approach: | They propose a framework that is free from referencing groundtruth annotations for investigating **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia's** on LLM and human judges. |
| Outcome: | The proposed framework investigates **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia' on LLM and human judges. |
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations (2025.emnlp-main)
Copied to clipboard
Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, null Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao
| Challenge: | Existing benchmarks assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments. |
| Approach: | They propose a new step-by-step evaluation pipeline that assesses formula selection, entity extraction, and arithmetic computation. |
| Outcome: | The proposed method improves the accuracy of large language models on medical benchmarks from 16.35% to 53.19%. |
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency (2025.emnlp-main)
Copied to clipboard
Svetlana Maslenkova, Clement Christophe, Marco AF Pimentel, Tathagata Raha, Muhammad Umar Salman, Ahmed Al Mahrooqi, Avani Gupta, Shadab Khan, Ronnie Rajan, Praveenkumar Kanithi
| Challenge: | Current dataset curation and bias assessment practices lack transparency . current approaches lack a thorough understanding of how data characteristics influence model behavior . |
| Approach: | They propose a comprehensive bias evaluation framework that integrates general benchmarks with a healthcare-specific methodology to probe for biases in a sensitive healthcare context. |
| Outcome: | The proposed approach to bias evaluation leverages established benchmarks and a healthcare-specific methodology. |
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown promising potential in the medical domain, assisting with tasks like clinical note generation and patient communication. |
| Approach: | They propose a framework that excels at utilizing domain-specific tools within two stages. |
| Outcome: | The proposed framework surpasses the pure LLMs with more than 10 points and the well-established agent-based methods with 3 points. |
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation (2026.eacl-demo)
Copied to clipboard
| Challenge: | EvalSense is a flexible framework for constructing domain-specific evaluation suites for large language models . it provides out-of-the-box support for a broad range of model providers and evaluation strategies . |
| Approach: | They propose a framework for constructing domain-specific evaluation suites for large language models. |
| Outcome: | The proposed framework provides out-of-the-box support for a broad range of model providers and evaluation strategies. |
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation (2025.acl-long)
Copied to clipboard
Xiechi Zhang, Zetian Ouyang, Linlin Wang, Gerard De Melo, Zhu Cao, Xiaoling Wang, Ya Zhang, Yanfeng Wang, Liang He
| Challenge: | Existing evaluation methods based on large language models (LLMs) are expensive and lack expertise due to limitations in human expertise. |
| Approach: | They propose an open-source automatic evaluation model with 13B parameters specifically engineered to measure the question-answering proficiency of medical LLMs. |
| Outcome: | The proposed model surpasses baselines in terms of correlation with human judgments. |