Papers by Xiongtao Zhou
MiCEval: Unveiling Multimodal Chain of Thought’s Quality via Image Description and Reasoning Steps (2025.naacl-long)
Copied to clipboard
Xiongtao Zhou, Jie He, Lanyu Chen, Jingyu Li, Haojing Chen, Victor Gutierrez Basulto, Jeff Z. Pan, Hanjie Chen
| Challenge: | Existing methods for evaluating the quality of reasoning steps in multimodal chain-of-thought are lacking. |
| Approach: | They propose a framework to evaluate the correctness of reasoning chains by evaluating the quality of both the description and each reasoning step. |
| Outcome: | The proposed framework improves interpretability and human judgments on four state-of-the-art MLLMs. |
An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Multimodal Large Language Models fine-tuned with multimodal instruction-following data have demonstrated formidable capabilities in multimodal tasks. |
| Approach: | They propose to employ four PEFT methods to fine-tune the LLM component of open-source MLLMs. |
| Outcome: | The proposed method is the best performing on seven datasets, while fine-tuning the connector layers leads to improved performance in most MLLMs. |