Papers by Jean-Benoit Delbrouck

7 papers
RadEval: A framework for radiology text evaluation (2025.emnlp-demos)

Copied to clipboard

Challenge: Evaluating automated radiology report generation systems remains a fundamental challenge in the development of safe, accurate, and clinically useful medical AI.
Approach: They propose a unified, open-source framework for evaluating radiology texts that consolidates a diverse range of metrics from classic ngram overlap (BLEU) and contextual measures (BERTScore) to clinical concept-based scores (GREEN).
Outcome: The framework consolidates a diverse range of metrics from ngram overlap (BLEU) and contextual measures (BERTScore) to clinical concept-based scores (F1CheXbert, F1RadGraph, RaTEScore, SRR-BERT, TemporalEntityF1) and advanced LLMbased evaluators (GREEN).
Improving the Factual Correctness of Radiology Report Generation with Semantic Rewards (2022.findings-emnlp)

Copied to clipboard

Challenge: Neural image-to-text radiology report generation systems have been successful on NLG metrics, but they are not factually complete or consistent due to inadequate training and evaluation.
Approach: They propose a method to improve the factual completeness and correctness of generated radiology reports by using a dataset containing annotated chest X-ray images.
Outcome: The proposed method significantly improves factual completeness and correctness of generated radiology reports on two open radiology report datasets.
GREEN: Generative Radiology Report Evaluation and Error Notation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing automated evaluation metrics fail to consider factual correctness or are limited in their interpretability.
Approach: They propose a radiology report evaluation metric that leverages natural language understanding of language models to identify and explain clinically significant errors.
Outcome: The proposed method demonstrates higher correlation with expert error counts and higher alignment with expert preferences when compared to previous methods.
RadGraph-XL: A Large-Scale Expert-Annotated Dataset for Entity and Relation Extraction from Radiology Reports (2024.findings-acl)

Copied to clipboard

Challenge: RadGraph-XL is an expert-annotated dataset for structured clinical data extraction.
Approach: They propose a large-scale, expert-annotated dataset for clinical entity and relation extraction using radiology reports.
Outcome: The proposed model outperforms existing methods by up to 52% and outperfies GPT-4 in this domain.
Toward Expanding the Scope of Radiology Report Summarization to Multiple Anatomies and Modalities (2023.acl-short)

Copied to clipboard

Challenge: Existing studies are limited to a single modality and a chest X-ray, making it difficult to replicate results or compare approaches.
Approach: They propose a dataset to generate an impression section of a radiology report . they propose to use three new modalities and seven new anatomies to evaluate their models .
Outcome: The proposed model is based on the MIMIC-III and MIMIC CXR datasets and evaluates their clinical efficacy via RadGraph, a factual correctness metric.
Automated Structured Radiology Report Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing models struggle to produce consistent, clinically meaningful reports and standard evaluation metrics fail to capture the nuances of radiological interpretation.
Approach: They propose to reformulate free-text radiology reports into a standardized format, ensuring clarity, consistency, and structured clinical reporting.
Outcome: The proposed task reformulates free-text radiology reports into a standardized format, ensuring clarity, consistency, and structured clinical reporting.
Structuring Radiology Reports: Challenging LLMs with Lightweight Models (2025.emnlp-main)

Copied to clipboard

Challenge: Radiology reports lack a standardized format, limiting both interpretability and machine learning applications.
Approach: They propose to use lightweight encoder-decoder models for structuring radiology reports . they compare models with eight open-source LLMs with prompting and in-context learning .
Outcome: The proposed models outperform eight open-source LLMs on a human-annotated test set.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations