Papers by Difei Gao
CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding (2023.acl-long)
Copied to clipboard
Zhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, W.k. Chan, Chong-Wah Ngo, Mike Zheng Shou, Nan Duan
| Challenge: | Existing work on video temporal grounding for long videos is limited by existing datasets. |
| Approach: | They propose a query-guided window selection strategy and a coarse-to-fine mechanism to speed up inference for long videos. |
| Outcome: | The proposed framework accelerates inference time by 2x on Ego4D-NLQ and 15x on MAD while keeping SOTA results. |
AssistSR: Task-oriented Video Segment Retrieval for Personal AI Assistant (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Currently, personal AI assistants on the phone and AR glasses can assist our daily life in addressing our questions like "how to adjust the date for this watch?" |
| Approach: | They propose a task that asks a question about affordance of items in our daily life . they construct a dataset that contains 3.2k multimodal questions on 1.6k video segments . |
| Outcome: | The proposed task outperforms baseline methods while still having room for improvement in the future. |
GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented Collaborations (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on the use of exocentric and egocentric videos in video question answering are focusing on eye-gaze information. |
| Approach: | They propose a task-oriented VQA dataset that captures eye-gaze information . they propose assisting models that ground the perceptual input into semantic information based on three different answer types . |
| Outcome: | The proposed model can ground the perceptual input into semantic information while reducing ambiguities. |