The More, The Better? A Critical Study of Multimodal Context in Radiology Report Summarization (2025.findings-emnlp)
Copied to clipboard
Mong Yuan Sim, Wei Emma Zhang, Xiang Dai, Biaoyan Fang, Sarbin Ranjitkar, Arjun Burlakoti, Jamie Taylor, Haojie Zhuang
| Challenge: | Current multimodal summarization models often fail to utilize radiology images in summarizing Findings section. |
| Approach: | They conduct a thorough analysis to determine whether current multimodal summarization models can utilize radiology images in summarizing Findings section. |
| Outcome: | The Impression section plays a crucial role in communication between radiologists and physicians. |
Similar Papers
Differentiable Multi-Agent Actor-Critic for Multi-Step Radiology Report Summarization (2022.acl-long)
Copied to clipboard
| Challenge: | Prior research on radiology report summarization has focused on single-step end-to-end models which subsume the task of salient content acquisition. |
| Approach: | They propose a two-step extractive summarization followed by abstractive summaries and a new method that breaks down the extractive part into two independent tasks: extraction of salient (1) sentences and (2) keywords. |
| Outcome: | The proposed model improves on English radiology reports with an overall improvement in F1 score of 3-4% compared to single-step and two-step-with-single-extractive-process baselines. |
Improving Radiology Summarization with Radiograph and Anatomy Prompts (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent studies focus on automatic impression generation, but this task is time-consuming and in high demand. |
| Approach: | They propose to use an anatomy-enhanced multimodal model to generate automatic impressions by combining radiology images with textual features. |
| Outcome: | The proposed model achieves state-of-the-art on two benchmark datasets and compares with existing models. |
Toward Expanding the Scope of Radiology Report Summarization to Multiple Anatomies and Modalities (2023.acl-short)
Copied to clipboard
| Challenge: | Existing studies are limited to a single modality and a chest X-ray, making it difficult to replicate results or compare approaches. |
| Approach: | They propose a dataset to generate an impression section of a radiology report . they propose to use three new modalities and seven new anatomies to evaluate their models . |
| Outcome: | The proposed model is based on the MIMIC-III and MIMIC CXR datasets and evaluates their clinical efficacy via RadGraph, a factual correctness metric. |
Multimodal Generation of Radiology Reports using Knowledge-Grounded Extraction of Entities and Relations (2022.aacl-main)
Copied to clipboard
Francesco Dalla Serra, William Clackett, Hamish MacKinnon, Chaoyang Wang, Fani Deligianni, Jeff Dalton, Alison Q. O’Neil
| Challenge: | Existing approaches to generate text radiology reports are prone to errors and poor clinical accuracy. |
| Approach: | They propose a two-step pipeline that subdivides the problem into factual triple extraction followed by free-text report generation. |
| Outcome: | The proposed pipeline shows that the generated reports exhibit realistic style but lack clinical accuracy. |
Word Graph Guided Summarization for Radiology Findings (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing studies focus on introducing salient word information to general text summarization framework to guide selection of key content in radiology findings. |
| Approach: | They propose a method for automatic impression generation using word graphs and a Word Graph guided Summarization model to capture critical words and their relations. |
| Outcome: | The proposed method is validated on two datasets, OPENI and MIMIC-CXR. |
Graph Enhanced Contrastive Learning for Radiology Findings Summarization (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods for automating impression generation have limited the relationship between extra knowledge and the original findings. |
| Approach: | They propose a framework for automating impression generation that exploits extra knowledge and original findings . they propose combining key words and their relations to extract critical information . |
| Outcome: | The proposed framework exploits extra knowledge and the original findings in an integrated way . the state-of-the-art results on two datasets confirm the effectiveness of the proposed method . |
From Sights to Insights: Towards Summarization of Multimodal Clinical Documents (2024.acl-long)
Copied to clipboard
| Challenge: | a recent WHO report highlights a drastic doctor-to-patient ratio . telehealth is one of the most impactful sectors where AI advances can bring a significant revolution . |
| Approach: | They propose an image-guided encoder-decoder model that uses contextual attention to create detailed visual-guides for multimodal documents. |
| Outcome: | The proposed model outperforms state-of-the-art models on multimodal question and dialogue summarization tasks. |
Show, Describe and Conclude: On Exploiting the Structure Information of Chest X-ray Reports (P19-1)
Copied to clipboard
| Challenge: | Existing studies do not consider the complex structure information between and within report sections. |
| Approach: | They propose a framework which exploits the structure information between and within report sections for generating CXR imaging reports. |
| Outcome: | The proposed framework achieves state-of-the-art performance on two CXR report datasets. |
A Dual-View Approach to Classifying Radiology Reports by Co-Training (2024.lrec-main)
Copied to clipboard
| Challenge: | Using the structure of a radiology report, we propose a co-training approach to train two machine learning models using the dual views of MRI and CT data. |
| Approach: | They propose a co-training approach where two machine learning models are built upon the Findings and Impression sections and use each other's information to boost performance with massive unlabeled data in a semi-supervised manner. |
| Outcome: | The proposed model outperforms supervised and semi-supervised methods in a public health surveillance study and outperformed existing methods. |
Pay More Attention to Images: Numerous Images-Oriented Multimodal Summarization (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing multimodal summarization approaches struggle with scenarios involving multiple images as input. |
| Approach: | They propose a task to generate multimodal summaries by integrating multiple images as input . they propose 'multimodal information evaluation' method that measures differences between generated summary and input based on multimodal input - and compares various methods . |
| Outcome: | The proposed method correlates more closely with human judgments than five widely used metrics . |