Papers by Jiatong Ma
M3-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering (2026.acl-long)
Copied to clipboard
| Challenge: | Existing knowledge-based VQA benchmarks focus on coarse-grained categories and simple reasoning over single entities. |
| Approach: | They propose a knowledge-based Visual Question Answering benchmark to enhance multimodality evaluation. |
| Outcome: | The proposed benchmark improves evaluation of multimodal large language models in fine-grained multimodal entity understanding and complex multihop reasoning. |
SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for multimodal large language models fail to capture complexity and traceability of reasoning processes . SciVQR includes domain-specific visuals and challenges models to combine visual comprehension with reasoning. |
| Approach: | They propose a multimodal benchmark for scientific reasoning covering 54 subfields . SciVQR includes domain-specific visuals and challenges models to combine visual comprehension with reasoning . |
| Outcome: | SciVQR evaluates 54 subfields in mathematics, physics, chemistry, geography, astronomy, and biology . the results highlight the need for improved multi-step reasoning and integration of interdisciplinary knowledge . |