Papers by Jialin Gao
Relation-aware Video Reading Comprehension for Temporal Language Grounding (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for temporal language grounding in videos are boundary regression and span extraction tasks. |
| Approach: | They propose a Relation-aware Network to localize a temporal span relevant to a given query sentence. |
| Outcome: | The proposed framework selects a video moment choice from the predefined answer set with the aid of coarse-and-fine choice-query interaction and choice-choice relation construction. |
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems (2025.findings-acl)
Copied to clipboard
| Challenge: | SciVerse is a multi-modal scientific evaluation benchmark to assess large multi-models . it examines the scientific knowledge comprehension, multi-mod content interpretation and Chain-of-Thought reasoning . authors examine the scientific proficiency of LMMs in scientific domains based on their work . |
| Approach: | They propose a multi-modal scientific evaluation benchmark to thoroughly assess Large Multi-modal Models across 5,735 test instances in five different versions. |
| Outcome: | The proposed evaluation reveals critical limitations in LMMs' scientific proficiency and provides new insights into future developments. |