| Challenge: | Existing methods for identifying human actions in videos are limited by the number of visual depictions in the videos. |
| Approach: | They propose a multimodal algorithm that leverages visual and linguistic clues to automatically infer which actions are visible in a video. |
| Outcome: | The proposed algorithm can identify actions visible in video while verbally describing them. |
Similar Papers
WhyAct: Identifying Action Reasons in Lifestyle Vlogs (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing systems for action recognition rely on pattern memorization and do not understand the action. |
| Approach: | They propose a multimodal model that leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video. |
| Outcome: | The proposed model leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video. |
ViLPAct: A Benchmark for Compositional Generalization on Multimodal Human Activities (2023.findings-eacl)
Copied to clipboard
Terry Yue Zhuo, Yaqing Liao, Yuecheng Lei, Lizhen Qu, Gerard de Melo, Xiaojun Chang, Yazhou Ren, Zenglin Xu
| Challenge: | a vision-language benchmark for human activity planning is designed for humans . the task is easy for humans, but challenging for SOTA deep learning models . |
| Approach: | They propose a vision-language benchmark for human activity planning that extends Charades with intents and builds on a multi-choice question test set. |
| Outcome: | The proposed benchmark evaluates the ability of systems to anticipate and plan human actions in a multimodal visionlanguage setting. |
Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows (2022.coling-1)
Copied to clipboard
Keisuke Shirai, Atsushi Hashimoto, Taichi Nishimura, Hirotaka Kameko, Shuhei Kurita, Yoshitaka Ushiku, Shinsuke Mori
| Challenge: | a new dataset enables us to learn a cooking action result for each object in a recipe text. |
| Approach: | They propose a multimodal dataset that enables us to learn a cooking action result for each object in a recipe text. |
| Outcome: | The proposed dataset reduces human annotation costs by allowing multimodal information retrieval. |
Intent Detection with WikiHow (2020.aacl-main)
Copied to clipboard
| Challenge: | Existing approaches to intent detection have limited data annotated for new domains or languages. |
| Approach: | They propose to train a set of pretraining intent detection models on wikiHow which can predict a broad range of intended goals from many actions. |
| Outcome: | The proposed models achieve state-of-the-art results on the Snips dataset, the Schema-Guided Dialogue dataset, and all 3 languages of the Facebook multilingual dialog datasets. |
LiMiT: The Literal Motion in Text Dataset (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Motion recognition is one of the basic cognitive capabilities of many life forms, yet identifying motion of physical entities in natural language have not been explored extensively and empirically. |
| Approach: | They propose to use a human-annotated dataset to identify motion of physical entities in natural language. |
| Outcome: | The proposed dataset analyzes the scale and diversity of the dataset and provides a baseline model. |
Connecting Language and Vision to Actions (P18-5)
Copied to clipboard
| Challenge: | Recent advances in language and vision have made incredible progress in describing images and interacting with visual content in a physical or embodied environment. |
| Approach: | This tutorial will provide an overview of the growing number of multimodal tasks and datasets that combine textual and visual understanding. |
| Outcome: | This tutorial will review the state-of-the-art approaches to selected tasks such as image captioning, visual question answering and visual dialog. |
Unveiling the Invisible: Captioning Videos with Metaphors (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have shown that Vision-Language models cannot understand visual metaphors in memes and adverts. |
| Approach: | They propose a task to describe visual metaphors in videos using a manually created dataset and a new metric called Average Concept Distance to automatically evaluate creativity. |
| Outcome: | The proposed system performs comparable to existing video language models on the proposed task and can be used for future research. |
Who is Speaking? Speaker-Aware Multiparty Dialogue Act Classification (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Identifying how speakers interact with each other in a conversation is difficult when more than two interlocutors take part in . To overcome this challenge, we propose to explicitly add speaker awareness to each utterance representation. |
| Approach: | They propose to add speaker awareness to each utterance representation to model how each speaker is behaving within the local context of a conversation. |
| Outcome: | The proposed approach is able to model multiparticipant and dyadic conversations on the MRDA and SwDA datasets and shows that it is more efficient than previous approaches. |
Recognizing Multimodal Entailment (2021.acl-tutorials)
Copied to clipboard
Cesar Ilharco, Afsaneh Shirazi, Arjun Gopalan, Arsha Nagrani, Blaz Bratanic, Chris Bregler, Christina Funk, Felipe Ferreira, Gabriel Barcik, Gabriel Ilharco, Georg Osang, Jannis Bulian, Jared Frank, Lucas Smaira, Qin Cao, Ricardo Marino, Roma Patel, Thomas Leung, Vaiva Imbrasaite
| Challenge: | This tutorial introduces the multimodal entailment task for detecting semantic alignments . the task requires fine-grained understanding of visual and linguistic semantics questions . |
| Approach: | This tutorial introduces the multimodal entailment task to machine learning . it introduces a dataset for recognizing multimodal alignments . |
| Outcome: | This tutorial introduces the multimodal entailment task . it can be useful for detecting semantic alignments when a single modality alone is not enough . |
Action Verb Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of 390 simple actions is based on multimodal data of 12 humans . the dataset is annotated with orthographic transcriptions of utterances and part-of-speech tags . |
| Approach: | They present a multimodal corpus of 12 humans performing 390 simple actions . they also propose an algorithm for segmenting words into utterances and aligning visual information and speech . |
| Outcome: | The presented dataset includes 390 simple actions performed by 12 humans . it includes transcriptions of utterances, part-of-speech tags, lemmata, and hand touches . |