Challenge: Existing methods for automatic Brain CT reports are limited by coarse-grained supervision and coupled cross-modal alignment.
Approach: They propose a pathological Graph-driven cross-modal alignment model that learns fine-grained visual cues and aligns them with textual words.
Outcome: The proposed model can improve the automatic generation of Brain CT reports and contribute to improved cranial disease diagnosis.

Similar Papers

See Detail Say Clear: Towards Brain CT Report Generation via Pathological Clue-driven Representation Learning (2024.findings-emnlp)

Copied to clipboard

Challenge: Brain CT report generation is important to aid physicians in diagnosing cranial diseases.
Approach: They propose a Pathological Clue-driven Representation Learning model to build cross-modal representations based on pathological clues and adapt them for text generation.
Outcome: The proposed method outperforms previous methods and achieves SoTA performance.
JPG - Jointly Learn to Align: Automated Disease Prediction and Radiology Report Generation (2022.coling-1)

Copied to clipboard

Challenge: Existing methods rarely consider cross-modal alignment between textual and visual features and ignore disease tags as auxiliary for report generation.
Approach: They propose a "Jointly learning framework for automated disease Prediction and radiology report Generation" the framework integrates cross-modal alignment between textual and visual features and disease tags to improve the quality of reports.
Outcome: The proposed framework improves the quality of radiology reports by combining the main task and auxiliary tasks.
KIA: Knowledge-Guided Implicit Vision-Language Alignment for Chest X-Ray Report Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing reports on medical images and reports lack fine-grained cross-modal interaction, leading to insufficient understanding of detailed information.
Approach: They propose a framework for establishing cross-modal semantic alignment in radiology report pairs using knowledge-guided implicit vision-language alignment.
Outcome: KIA improves understanding of medical images and reports by incorporating medical knowledge to enhance pathological observation and anatomical landm.
CmEAA: Cross-modal Enhancement and Alignment Adapter for Radiology Report Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for automatic radiology report generation suffer from data bias.
Approach: They propose a method that connects a vision encoder with a frozen large language model by using a cross-modal enhancement and alignment adapter.
Outcome: The proposed model outperforms existing state-of-the-art methods on IU X-Ray and MIMIC-CXR datasets.
Reinforced Cross-modal Alignment for Radiology Report Generation (2022.findings-acl)

Copied to clipboard

Challenge: Medical images are widely used in clinical decision-making, where writing radiology reports can be enhanced by automatic solutions to alleviate physicians’ workload.
Approach: They propose an approach with reinforcement learning over a cross-modal memory to better align visual and textual features for radiology report generation.
Outcome: The proposed approach improves cross-modal alignment on two English radiology report datasets and human evaluation confirms the results.
Cross-modal Contrastive Attention Model for Medical Report Generation (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for medical report generation are unable to capture useful information from historical cases.
Approach: They propose a model that captures both visual and semantic information from similar cases.
Outcome: The proposed model outperforms the state-of-the-art models on almost all metrics on IU X-Ray and MIMIC-CXR benchmarks.
Fine-grained Medical Vision-Language Representation Learning for Radiology Report Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to learn medical vision-language representations by contrasting images with entire reports are not effective.
Approach: They propose a phenotype-driven medical vision-language representation learning framework to bridge the gap between visual and textual modalities for improved text-oriented generation.
Outcome: The proposed framework bridges the gap between visual and textual modalities for improved radiology report generation.
Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework (2026.acl-long)

Copied to clipboard

Challenge: Current methods map whole volumes to reports, ignoring the clinical workflow of analyzing localized Regions of Interest (RoIs) Current models exhibit suboptimal accuracy and are prone to significant hallucinations.
Approach: They propose a framework that mimics the professional radiologist diagnostic workflow by employing graph-based relational modules to capture dependencies between RoI attributes.
Outcome: The proposed framework surpasses existing models by 19.7% in BLEU and 4.7% in ROUGE-L while achieving a 45.8% improvement in clinical metrics.
Learning Visual-Semantic Embeddings for Reporting Abnormal Findings on Chest X-rays (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing work on report generation often trains encoder-decoder networks to generate complete reports, but such models are affected by data bias and face common issues inherent in text generation models.
Approach: They propose a method to identify abnormal findings from radiology images and group them with unsupervised clustering and minimal rules.
Outcome: The proposed method outperforms existing generation models on correctness and text generation metrics.
X-ray Made Simple: Lay Radiology Report Generation and Robust Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Technical language and templated nature of professional reports hinder patient comprehension and allow models to artificially boost lexical metrics such as BLEU by reproducing common report patterns.
Approach: They propose a layman's RRG framework that leverages layperson-friendly language to enhance patient accessibility and promote robust evaluation and report generation by encouraging models to focus on semantic accuracy over rigid templates.
Outcome: The proposed framework improves model performance with more layman-style data, compared to templated professional language and inflated lexical scores.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations