Ivan Sviridov, Amina Miftakhova, Tereshchenko Artemiy Vladimirovich, Galina Zubkova, Pavel Blinov, Andrey Savchenko
| Challenge: | Large Vision-Language Models (LVLMs) are being explored in medicine but their ability to conduct complex real-world telemedicine consultations remains underexplored. |
| Approach: | They propose to use large vision-language models to conduct telemedicine consultations using a framework that simulates patient variability and evaluates diagnostic accuracy and dialogue quality via Assessor Agent. |
| Outcome: | The proposed framework compares diagnostic strategies for open and closed-source LVLMs and shows that multimodal dialogue improves F1 score by 6.5% over non-dialogue settings. |
Similar Papers
Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image Understanding (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for multimodal large language models are limited to multiview diagnostics. |
| Approach: | They propose a benchmark specifically designed for medical multi-image understanding that evaluates MLLMs across four dimensions. |
| Outcome: | The proposed model performs better in multi-image contexts than open-source models . the model perform better when processing increased visual loads than closed-source ones . |
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on basic abilities using nonverbal methods, such as yes-no and multiple-choice questions. |
| Approach: | They propose a benchmark that provides more nuanced evaluations of alignment capabilities for large Vision-Language Models (VLMs) they use a rule-calibrated evaluator that exceeds GPT-4's evaluation ability and a “alignment score” to assess the robustness and stability of models across diverse prompts. |
| Outcome: | The proposed benchmark covers 13 tasks across three categories and includes both single-turn and multi-turn dialogue scenarios. |
LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models (2025.findings-naacl)
Copied to clipboard
Zhenyue Qin, Yu Yin, Dylan Campbell, Xuansheng Wu, Ke Zou, Ninghao Liu, Yih Chung Tham, Xiuzhen Zhang, Qingyu Chen
| Challenge: | Existing benchmarks for large vision-language models (LVLMs) are limited to ophthalmology-specific applications. |
| Approach: | They introduce a large-scale multimodal ophthalmology benchmark consisting of 21,993 instances across five ocular imaging modalities and 13 state-of-the-art LVLM representatives from closed-source, open-source and medical domains. |
| Outcome: | The proposed model shows significant performance drop in ophthalmology compared to other domains. |
MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Medical Vision-Language Models are predominantly trained on professional literature, limiting their ability to communicate findings in the lay register required for patient-centered care. |
| Approach: | They propose a multimodal benchmark dedicated to expert-lay semantic alignment that enforces strict semantic equivalence by integrating unified medical language system (UMS) Concept Unique Identifiers (CUIs) with micro-level entity constraints. |
| Outcome: | The proposed benchmark enforces strict semantic equivalence by integrating unified medical language system (UMLS) Concept Unique Identifiers (CUIs) with micro-level entity constraints. |
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation (2025.findings-naacl)
Copied to clipboard
João Matos, Shan Chen, Siena Kathleen V. Placino, Yingya Li, Juan Carlos Climent Pardo, Daphna Idan, Takeshi Tohyama, David Restrepo, Luis Filipe Nakayama, José María Millet Pascual-Leone, Guergana K Savova, Hugo Aerts, Leo Anthony Celi, An-Kwok Ian Wong, Danielle Bitterman, Jack Gallifant
| Challenge: | Existing multiple-choice question and answer (QA) datasets are text-only and available in a limited subset of languages and countries. |
| Approach: | They propose a multilingual, multimodal benchmarking dataset to evaluate multimodal/vision language models in healthcare. |
| Outcome: | The WorldMedQA-V includes 568 labeled multiple-choice QAs paired with 568 medical images from four countries. |
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues (2026.findings-acl)
Copied to clipboard
Yaning Pan, Qianqian Xie, Guohui Zhang, Zekun Moore Wang, Yongqian Wen, Yuanxing Zhang, Haoxuan Hu, Zhiyu Pan, Yibing Huang, Zhidong Gan, Yonghong Lin, An Ping, Shihao Li, Yanghai Wang, Tianhao Peng, Jiaheng Liu
| Challenge: | Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios. |
| Approach: | They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity. |
| Outcome: | The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues. |
MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions (2025.emnlp-main)
Copied to clipboard
Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha
| Challenge: | omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants. |
| Approach: | They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models. |
| Outcome: | The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. |
WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for large language models focus on webpage generation outcomes. |
| Approach: | They propose a multi-view evaluation framework to evaluate MLLMs in four key areas: WebUI Perception, HTML Programming, WebUI-HTML Understanding, and WebUI to code. |
| Outcome: | The proposed framework evaluates MLLMs in four key areas: WebUI Perception, HTML Programming, WebUI-HTML Understanding, and WebUI to code. |
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation (2026.acl-long)
Copied to clipboard
Yanghao Zhou, Haitian Li, Rexar Lin, Heyan Huang, Jinxing Zhou, Changsen Yuan, Tian Lan, Ziqin Zhou, Yudong Li, Jiajun Xu, Jingyun Liao, YiMing Cheng, Xuefeng Chen, Xian-Ling Mao, Yousheng Feng
| Challenge: | Existing evaluation benchmarks for text-to-audio-video (T2AV) generation are largely designed for human-recorded videos or single-speaker settings. |
| Approach: | They propose a failure-driven diagnostic benchmark for multi-talker dialogue-centric audio-video generation. |
| Outcome: | The benchmark evaluates multi-speaker dialogue generation at four levels: audio-visual signal fidelity, temporal attribute consistency, social interaction, and cinematic expression. |
MEDSYN: Benchmarking Multi-EviDence SYNthesis in Complex Clinical Cases for Multimodal Large Language Models (2026.findings-acl)
Copied to clipboard
Boqi Chen, Xudong Liu, Jiachuan Peng, Marianne Frey-Marti, Kyle Lam, Bang Zheng, Lin Li, Jianing Qiu
| Challenge: | Existing benchmarks for multimodal large language models do not capture real-world clinical complexity. |
| Approach: | They evaluate multilingual, multimodal multimodal models of clinical cases with up to 7 distinct visual clinical evidence types per case. |
| Outcome: | The proposed model outperforms human models on differential diagnosis (DDx) generation and final diagnosis (FDx) selection. |