Challenge: Visual Language Models (VLMs) have shown strong performance in tasks like radiology report generation but struggle with hallucinations, vague descriptions, Inconsistent logic and poor localization.
Approach: They propose a framework for medical visual reasoning based on Visual Guidance and Self-Reward paradigms and Monte Carlo Tree Search to improve the model's visual reasoning capabilities.
Outcome: The proposed framework outperforms existing models on multiple medical VQA benchmarks.

Similar Papers

MedThink: A Rationale-Guided Framework for Explaining Medical Visual Question Answering (2025.findings-naacl)

Copied to clipboard

Challenge: Existing models for medical visual question answering are limited in their interpretation and interpretation . a semi-automated annotation process is used to streamline data preparation and build new benchmark datasets .
Approach: They propose a semi-automated annotation process to streamline data preparation and build new benchmark Med-VQA datasets.
Outcome: The proposed method achieves an accuracy of 83.5% on R-RAD, 86.3% on RSLAKE and 87.2% on RPath.
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Large language models face unique challenges such as domain-specific terminologies and reasoning over specialized knowledge.
Approach: They propose a multi-disciplinary collaboration framework that leverages LLM-based agents in a role-playing setting.
Outcome: The proposed framework excels at mining and harnessing medical expertise within LLMs, as well as extending its reasoning abilities.
Guiding Medical Vision-Language Models with Diverse Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations (2025.naacl-long)

Copied to clipboard

Challenge: Current vision-language models lack the ability to focus on specific areas designated by humans . a new framework that integrates medical entity extraction, visual prompt generation, and dataset adaptation is proposed to improve visual prompt-guided fine-tuning.
Approach: They propose to use visual prompts to guide and enhance formation of region-specific attention.
Outcome: The proposed framework outperforms state-of-the-art large vision-language models on medical datasets.
Med-SRAF: A Multi-Agent Framework for Medical Reasoning via Semantic Routing and Agentic Fusion (2026.findings-acl)

Copied to clipboard

Challenge: Existing RAG methods suffer from a two-part problem: semantic drift and concatenation fallacy . et al.: rapid development of Large Language Models has led to a paradigm shift in artificial intelligence .
Approach: They propose a multi-agent retrieval augmentation framework guided by medical domain knowledge to address these challenges.
Outcome: The proposed framework outperforms existing general RAG baselines on five widely used medical benchmarks.
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts.
Approach: They propose a novel agentic framework that explicitly performs visual reasoning directly within the chart’s spatial domain.
Outcome: The proposed framework achieves state-of-the-art accuracy on the ChartBench and ChartX benchmarks surpassing prior methods by up to 16.07% absolute gain overall and 17.31% on numerically intensive queries.
AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing Med-MLLMs fail when deployed in low-resource settings where abundant labeled data is unavailable.
Approach: They propose a training-free agentic framework that performs medical knowledge augmentation via LLM agents.
Outcome: The proposed framework performs medical knowledge augmentation via LLM agents.
ViLMedic: a framework for research at the intersection of vision and language in medical AI (2022.acl-demo)

Copied to clipboard

Challenge: Multimodal medical AI is a growing field of interest, especially for tasks that involve multimodal data.
Approach: They propose a vision-and-language medical library to improve multimodal medical predictions and enable new applications.
Outcome: The vision-and-language medical library aims to improve reproducibility and speed up progress across medical AI . it contains a dozen implementations replicating state-of-the-art results on medical datasets . the library is extensible by researchers but also simple for practitioners .
A Survey of LLM-based Agents in Medicine: How far are we from Baymax? (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming healthcare through their ability to understand and assist with medical tasks.
Approach: They analyze system profiles, clinical planning, medical reasoning frameworks, and external capacity enhancement.
Outcome: The findings highlight the future directions in medical reasoning, physical system integration, and training simulations.
Beyond Surface Features: Advancing Medical Vision-Language Alignment via Dynamic Evidence-Guided Preference Optimization (2026.acl-long)

Copied to clipboard

Challenge: Existing preference-based methods for medical large vision-Language Models face limitations in medical settings . existing methods are limited by overfitting to superficial cues and pseudo convergence of the preference signal.
Approach: They propose a framework that enables evidence-aware and adaptive preference learning for Med-LVLMs.
Outcome: The proposed framework improves evidence-aware and adaptive preference learning for Med-LVLMs.
MedCoT: Medical Chain of Thought via Hierarchical Expert (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for medical visual question answering lack robustness and reasoning paths for real-world medical diagnostics.
Approach: They propose a hierarchical expert verification reasoning chain method to enhance interpretability and accuracy in medical visual question answering.
Outcome: The proposed method outperforms existing methods on four standard Med-VQA datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations