| Challenge: | Existing systems for action recognition rely on pattern memorization and do not understand the action. |
| Approach: | They propose a multimodal model that leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video. |
| Outcome: | The proposed model leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video. |
Similar Papers
Identifying Visible Actions in Lifestyle Vlogs (P19-1)
Copied to clipboard
| Challenge: | Existing methods for identifying human actions in videos are limited by the number of visual depictions in the videos. |
| Approach: | They propose a multimodal algorithm that leverages visual and linguistic clues to automatically infer which actions are visible in a video. |
| Outcome: | The proposed algorithm can identify actions visible in video while verbally describing them. |
ViLPAct: A Benchmark for Compositional Generalization on Multimodal Human Activities (2023.findings-eacl)
Copied to clipboard
Terry Yue Zhuo, Yaqing Liao, Yuecheng Lei, Lizhen Qu, Gerard de Melo, Xiaojun Chang, Yazhou Ren, Zenglin Xu
| Challenge: | a vision-language benchmark for human activity planning is designed for humans . the task is easy for humans, but challenging for SOTA deep learning models . |
| Approach: | They propose a vision-language benchmark for human activity planning that extends Charades with intents and builds on a multi-choice question test set. |
| Outcome: | The proposed benchmark evaluates the ability of systems to anticipate and plan human actions in a multimodal visionlanguage setting. |
Everything Happens for a Reason: Discovering the Purpose of Actions in Procedural Text (D19-1)
Copied to clipboard
| Challenge: | XPAD is a new model that predicts actions' effects and their dependencies based on background knowledge . previous work on extracting sequences of actions from text has focused on identifying why they are the way they are . |
| Approach: | They propose a new model that biases effect predictions towards those that explain more of the actions in the paragraph and are more plausible with respect to background knowledge. |
| Outcome: | The proposed model outperforms existing systems on explaining actions by predicting dependencies while maintaining the performance on the original task in ProPara. |
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation (2026.acl-long)
Copied to clipboard
Ziyi Wang, Yuxuan Lu, Wenbo Li, Amirali Amini, Bo Sun, Yakov Bart, Weimin Lyu, Jiri Gesi, Tian Wang, Jing Huang, Yu Su, Upol Ehsan, Malihe Alikhani, Toby Jia-Jun Li, Lydia Chilton, Dakuo Wang
| Challenge: | evaluating LLMs' ability to mimic real user behavior remains an open challenge due to the lack of high-quality, publicly available datasets that capture both the observable actions and the internal reasoning of an actual user. |
| Approach: | They propose a dataset of Observation, Persona, Rationale, and Action collected from real human participants during online shopping sessions. |
| Outcome: | The proposed dataset is the first to evaluate how well current LLMs can accurately simulate the next web action of a specific user. |
JX4MEI: Multimodal Semantically-Enhanced LLM for Joint Multimodal Emotion-Intent Explanation and Classification (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing multimodal emotion and intent recognition tasks focus on classification, not rationale and intrinsic connections between these states. |
| Approach: | They propose a task that requires models to jointly predict emotion and intent while generating natural language explanations for why they co-occur. |
| Outcome: | The proposed model outperforms baseline models in prediction and explanation generation. |
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | Despite advances in artificial intelligence, building social intelligence remains a challenge. |
| Approach: | They propose a task to explain why people laugh in a video and a dataset to do this. |
| Outcome: | The proposed dataset generates plausible explanations for laughter in video and in-the-wild videos. |
Predicting Human Activities from User-Generated Content (P19-1)
Copied to clipboard
| Challenge: | Several studies have applied computational approaches to the understanding and modeling of human behavior at scale and in real time. |
| Approach: | They propose a sentence embedding framework tailored to recognize the semantics of human activities and perform automatic clustering of these activities. |
| Outcome: | The proposed framework can make predictions based on the text of user-generated content and self-description. |
Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies focus on utterance-level emotion cause extraction and multimodal emotion cause generation, resulting in subjective and inconsistent annotations. |
| Approach: | They propose a task that extracts emotion cause utterances and generates cause summaries . they propose utterrance-level emotion cause extraction and multimodal emotion cause generation tasks . |
| Outcome: | The proposed task extracts emotion cause utterances and generates cause summaries . the proposed task establishes strong benchmark results for the proposed project . |
TellMeWhy: A Dataset for Answering Why-Questions in Narratives (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing models do not have the ability to answer "why" questions that require commonsense knowledge external to the narrative. |
| Approach: | They propose a crowd-sourced dataset that asks why characters perform actions . they show that state-of-the-art models are far below human performance on answering such questions . |
| Outcome: | The proposed dataset shows that state-of-the-art models are far below human performance on answering such questions. |
Causal Explanation Analysis on Social Media (D18-1)
Copied to clipboard
| Challenge: | Understanding causal explanations is an important psychological factor linked to physical and mental health. |
| Approach: | They propose to automate causal explanation analysis by building on discourse parsing and using a hierarchy of Bidirectional LSTMs to identify the specific phrase that is the explanation. |
| Outcome: | The proposed subtasks achieve strong accuracies but differ in their approaches . the proposed sub task is compared with the previous task and is able to identify the specific phrase that is the explanation. |