Papers by Hanwen Miao
WhyAct: Identifying Action Reasons in Lifestyle Vlogs (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing systems for action recognition rely on pattern memorization and do not understand the action. |
| Approach: | They propose a multimodal model that leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video. |
| Outcome: | The proposed model leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video. |