Papers with Med-LVLMs

5 papers
RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing medical large vision language models often generate inaccurate and irrelevant answers that do not align with established medical facts.
Approach: They propose a strategy for controlling factuality risk through calibrated selection of the number of retrieved contexts and a preference dataset to fine-tune the model.
Outcome: The proposed model achieves an average improvement of 20.8% on three medical VQA datasets.
HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks (2026.findings-acl)

Copied to clipboard

Challenge: Medical large vision-language models suffer from factual inaccuracies and unreliable outputs.
Approach: They propose a framework that enhances Med-LVLMs through heterogeneous knowledge sources.
Outcome: The proposed framework improves Med-LVLMs through heterogeneous knowledge sources.
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods rely on inference-time interventions, which are limited in attention adaptation or require additional supervision.
Approach: They propose a framework for automatic attention alignment tuning that leverages weak labels from SAM and selectively modifies visually-critical attention heads to improve alignment while minimizing interference.
Outcome: The proposed framework outperforms state-of-the-art models on medical VQA and report generation benchmarks.
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback (2025.acl-long)

Copied to clipboard

Challenge: Existing Medical Large Vision-Language Models (Med-LVLMs) lack visual localization in medical images, which is crucial for abnormality detection and interpretation.
Approach: They propose a medical abnormalities unveiling method based on a Medical Abnormalities Unveiler dataset and propose 'abnormal-aware instruction tuning' and 'abbnormal-Aware Reward' method generates diagnoses based upon identified abnormal areas in medical images.
Outcome: The proposed method outperforms existing medical large vision-language models in identifying and understanding medical abnormalities and improves generalization capability.
Beyond Surface Features: Advancing Medical Vision-Language Alignment via Dynamic Evidence-Guided Preference Optimization (2026.acl-long)

Copied to clipboard

Challenge: Existing preference-based methods for medical large vision-Language Models face limitations in medical settings . existing methods are limited by overfitting to superficial cues and pseudo convergence of the preference signal.
Approach: They propose a framework that enables evidence-aware and adaptive preference learning for Med-LVLMs.
Outcome: The proposed framework improves evidence-aware and adaptive preference learning for Med-LVLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations