Papers by Yinan Zou
Can Multimodal Large Language Models Understand Spatial Relations? (2025.acl-long)
Copied to clipboard
| Challenge: | Spatial relation reasoning is a crucial task for multimodal large language models to understand the objective world. |
| Approach: | They propose a human-annotated spatial relation reasoning benchmark based on COCO2017 to improve MLLMs' spatial relation thinking. |
| Outcome: | The proposed benchmark achieves 48.14% accuracy, far below the human-level accuracy of 98.40%. |