Papers with MSVD-QA
Multi-Scale Progressive Attention Network for Video Question Answering (2021.acl-short)
Copied to clipboard
| Challenge: | Experimental evaluations on three benchmarks: TGIF-QA, MSVD-QA and MSRVTT-QA show our method has achieved state-of-the-art performance. |
| Approach: | They propose a multi-scale progressive attention network to fuse visual and text information. |
| Outcome: | The proposed method achieves state-of-the-art on three benchmarks: TGIF-QA, MSVD-QA and MSRVTT-QA. |
VideoQA-TA: Temporal-Aware Multi-Modal Video Question Answering (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for video question answering align visual or textual features directly with large language models, limiting the deep semantic association between modalities and hindering a comprehensive understanding of interactions within spatial and temporal contexts. |
| Approach: | They propose a temporal-aware framework for multi-modal video question answering that aligns videos and questions at fine-grained levels. |
| Outcome: | The proposed framework improves reasoning ability and accuracy of videoQA by aligning videos and questions at fine-grained levels. |