Challenge: a large part of human communication relies on nonverbal cues such as facial expressions, eye contact, and body language.
Approach: They propose to validate whether video large language models can correctly interpret body language from short clips of body language.
Outcome: The proposed model can correctly interpret emotions from short clips of body language.

Similar Papers

POSQA: Probe the World Models of LLMs with Size Comparisons (2023.findings-emnlp)

Copied to clipboard

Challenge: Embodied language comprehension emphasizes that language understanding is not only mental processing in the brain but also involves interactions with the physical and social environment.
Approach: They propose to use a physical object size question to examine the extremity of large language models to test their embodied comprehension.
Outcome: The proposed dataset shows that even the largest LLMs perform poorly under the zero-shot setting.
Evaluating Large Vision Language Models on Bangla Medical Visual Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models and Large Vision Language Model (LVLMs) have demonstrated promising capabilities in complex reasoning tasks, but low-resource contexts like Bangla are underexplored.
Approach: They propose a multilingual medical visual question answering dataset using Bangla.
Outcome: The proposed model performs well on generalized visual tasks but struggles with fine-grained diagnostic reasoning, achieving low accuracy in specialized categories.
Anatomy of a Feeling: Narrating Embodied Emotions via Large Vision-Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: ELENA is a framework for embodied emotion analysis using large vision language models . ELEna uses attention maps and a persistent bias towards the facial region .
Approach: They propose a framework that utilizes large vision language models to generate ELENA . they propose to use attention maps to describe emotional reactions from body parts .
Outcome: The proposed framework outperforms baseline models without fine-tuning . it uses large vision language models to generate embodied emotion narratives .
VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions (2025.acl-long)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) can grasp the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent.
Approach: They propose a framework for multimodal large language models to grasp the intention of a question and decompose it into a series of visual recognition sub-tasks to find out the answer.
Outcome: The proposed framework improves the accuracy of complex video-related questions by 29.6% and 17.2% on CVQA and the existing VQA datasets.
Do Video Language Models really understand the video contexts? (2025.naacl-srw)

Copied to clipboard

Challenge: Recent advances in VideoQA performance have shown that visual language models are effective but the processes of understanding and reasoning in VLMs remain under-explored.
Approach: They propose a framework that incorporates a fine-grained question generation and answering process to measure how well VLMs understand video question answering tasks.
Outcome: The proposed framework incorporates a fine-grained question generation and answering process to measure how well the responses generated by VLMs align with what the model understands.
LifeQA: A Real-life Dataset for Video Question Answering (2020.lrec-1)

Copied to clipboard

Challenge: Existing video question answering datasets consist of movies and TV shows, but they are not representative of our day-to-day lives.
Approach: They propose a benchmark dataset for video question answering that focuses on day-to-day situations.
Outcome: The proposed dataset analyzes the challenging but realistic aspects of LifeQA . it consists of video clips and over 2.3k multiple-choice questions .
Large Language Models are Temporal and Causal Reasoners for Video Question Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable performances on a wide range of natural language understanding and generation tasks.
Approach: They propose a framework that exploits linguistic shortcuts and mitigates 'linguistic bias' by flipping the source pair and target label to understand their complex relationships.
Outcome: The proposed framework outperforms both LLMs-based and non-LLMs- based models on five challenging VideoQA benchmarks.
In-the-Wild Video Question Answering (2022.coling-1)

Copied to clipboard

Challenge: Existing video understanding datasets focus on human interactions with little attention being paid to the “in the wild” settings.
Approach: They propose a video understanding dataset of videos recorded outdoors . they propose identifying visual support for a given question and answer .
Outcome: The proposed dataset examines the ability of models to understand videos, including video question answering, video captioning, and fill-inthe-blank tasks.
NegVQA: Can Vision Language Models Understand Negation? (2025.findings-acl)

Copied to clipboard

Challenge: NegVQA is a visual question answering (VQA) benchmark consisting of 7,379 two-choice questions covering diverse negation scenarios and image-question distributions.
Approach: They propose a visual question answering benchmark consisting of 7,379 two-choice questions covering diverse negation scenarios and image-question distributions.
Outcome: The proposed model fails to correctly interpret negation, leading to critical errors in interactive AI systems.
A Unified View on Emotion Representation in Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Recent studies show the presence of emotion concepts in the hidden state representations, but it’s unclear if the model has a robust representation consistent across different datasets.
Approach: They propose a unified view to understand emotion representation in Large Language Models by experimenting with diverse datasets and prompts.
Outcome: The proposed model can be interchanged between datasets with minimal impact on performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations