Papers by Chuyuan Wang
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning (2025.coling-main)
Copied to clipboard
Yifan Du, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, Jinpeng Wang, Chuyuan Wang, Mingchen Cai, Ruihua Song, Ji-Rong Wen
| Challenge: | Experimental results show that visual instruction tuning improves performance of Multi-modal Large Language Models (MLLMs) to extend the application scope of Large Language Modells, a surge of work augments LLMs with vision encoders to endow the ability of multi-modal cognition and reasoning. |
| Approach: | They propose a systematic approach to create high-quality visual reasoning instructions using a synthesize-complicate-reformulate paradigm. |
| Outcome: | The proposed method improves performance of MLLMs by 27.86% and 27.60% on MME-Perception and MME Cognition. |