Challenge: Current multimodal summarization models often fail to utilize radiology images in summarizing Findings section.
Approach: They conduct a thorough analysis to determine whether current multimodal summarization models can utilize radiology images in summarizing Findings section.
Outcome: The Impression section plays a crucial role in communication between radiologists and physicians.

Similar Papers

Differentiable Multi-Agent Actor-Critic for Multi-Step Radiology Report Summarization (2022.acl-long)

Copied to clipboard

Challenge: Prior research on radiology report summarization has focused on single-step end-to-end models which subsume the task of salient content acquisition.
Approach: They propose a two-step extractive summarization followed by abstractive summaries and a new method that breaks down the extractive part into two independent tasks: extraction of salient (1) sentences and (2) keywords.
Outcome: The proposed model improves on English radiology reports with an overall improvement in F1 score of 3-4% compared to single-step and two-step-with-single-extractive-process baselines.
Improving Radiology Summarization with Radiograph and Anatomy Prompts (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies focus on automatic impression generation, but this task is time-consuming and in high demand.
Approach: They propose to use an anatomy-enhanced multimodal model to generate automatic impressions by combining radiology images with textual features.
Outcome: The proposed model achieves state-of-the-art on two benchmark datasets and compares with existing models.
Toward Expanding the Scope of Radiology Report Summarization to Multiple Anatomies and Modalities (2023.acl-short)

Copied to clipboard

Challenge: Existing studies are limited to a single modality and a chest X-ray, making it difficult to replicate results or compare approaches.
Approach: They propose a dataset to generate an impression section of a radiology report . they propose to use three new modalities and seven new anatomies to evaluate their models .
Outcome: The proposed model is based on the MIMIC-III and MIMIC CXR datasets and evaluates their clinical efficacy via RadGraph, a factual correctness metric.
Multimodal Generation of Radiology Reports using Knowledge-Grounded Extraction of Entities and Relations (2022.aacl-main)

Copied to clipboard

Challenge: Existing approaches to generate text radiology reports are prone to errors and poor clinical accuracy.
Approach: They propose a two-step pipeline that subdivides the problem into factual triple extraction followed by free-text report generation.
Outcome: The proposed pipeline shows that the generated reports exhibit realistic style but lack clinical accuracy.
Word Graph Guided Summarization for Radiology Findings (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on introducing salient word information to general text summarization framework to guide selection of key content in radiology findings.
Approach: They propose a method for automatic impression generation using word graphs and a Word Graph guided Summarization model to capture critical words and their relations.
Outcome: The proposed method is validated on two datasets, OPENI and MIMIC-CXR.
Graph Enhanced Contrastive Learning for Radiology Findings Summarization (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for automating impression generation have limited the relationship between extra knowledge and the original findings.
Approach: They propose a framework for automating impression generation that exploits extra knowledge and original findings . they propose combining key words and their relations to extract critical information .
Outcome: The proposed framework exploits extra knowledge and the original findings in an integrated way . the state-of-the-art results on two datasets confirm the effectiveness of the proposed method .
From Sights to Insights: Towards Summarization of Multimodal Clinical Documents (2024.acl-long)

Copied to clipboard

Challenge: a recent WHO report highlights a drastic doctor-to-patient ratio . telehealth is one of the most impactful sectors where AI advances can bring a significant revolution .
Approach: They propose an image-guided encoder-decoder model that uses contextual attention to create detailed visual-guides for multimodal documents.
Outcome: The proposed model outperforms state-of-the-art models on multimodal question and dialogue summarization tasks.
Show, Describe and Conclude: On Exploiting the Structure Information of Chest X-ray Reports (P19-1)

Copied to clipboard

Challenge: Existing studies do not consider the complex structure information between and within report sections.
Approach: They propose a framework which exploits the structure information between and within report sections for generating CXR imaging reports.
Outcome: The proposed framework achieves state-of-the-art performance on two CXR report datasets.
A Dual-View Approach to Classifying Radiology Reports by Co-Training (2024.lrec-main)

Copied to clipboard

Challenge: Using the structure of a radiology report, we propose a co-training approach to train two machine learning models using the dual views of MRI and CT data.
Approach: They propose a co-training approach where two machine learning models are built upon the Findings and Impression sections and use each other's information to boost performance with massive unlabeled data in a semi-supervised manner.
Outcome: The proposed model outperforms supervised and semi-supervised methods in a public health surveillance study and outperformed existing methods.
Pay More Attention to Images: Numerous Images-Oriented Multimodal Summarization (2025.naacl-long)

Copied to clipboard

Challenge: Existing multimodal summarization approaches struggle with scenarios involving multiple images as input.
Approach: They propose a task to generate multimodal summaries by integrating multiple images as input . they propose 'multimodal information evaluation' method that measures differences between generated summary and input based on multimodal input - and compares various methods .
Outcome: The proposed method correlates more closely with human judgments than five widely used metrics .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations