Papers by Gi-Cheon Kang
Attend What You Need: Motion-Appearance Synergistic Networks for Video Question Answering (2021.acl-long)
Copied to clipboard
| Challenge: | Recent advances in natural language processing and computer vision have made significant progress in artificial intelligence (AI). |
| Approach: | They propose Motion-Appearance Synergistic Networks which embed cross-modal features grounded on motion and appearance information and selectively utilize them depending on the question’s intentions. |
| Outcome: | The proposed network achieves state-of-the-art on the TGIF-QA and MSVD-QA datasets and qualitatively analyzes the results. |
Dual Attention Networks for Visual Reference Resolution in Visual Dialog (D19-1)
Copied to clipboard
| Challenge: | Visual dialog (VisDial) requires a dialog agent to answer a series of questions grounded in an image. |
| Approach: | They propose dual attention networks (DAN) for visual reference resolution in VisDial. |
| Outcome: | The proposed model outperforms the previous state-of-the-art model on VisDial datasets. |
Reasoning Visual Dialog with Sparse Graph Learning and Knowledge Transfer (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Visual dialog is a task of answering questions grounded in an image using dialog history as context. |
| Approach: | They propose a Sparse Graph Learning method to formulate visual dialog as a graph structure learning task. |
| Outcome: | The proposed model outperforms the state-of-the-art models on the VisDial v1.0 dataset. |