Challenge: Humor enriches our daily lives and appears in many forms, from jokes and cartoons to comedies and viral videos.
Approach: They introduce a video humor understanding benchmark to test their ability to understand humor from visual cues.
Outcome: The proposed video humor understanding benchmark is based on a collection of short videos . it features rich annotations and a study of environmental sound that can enhance humor .

Similar Papers

Can Language Models Laugh at YouTube Short-form Videos? (2023.emnlp-main)

Copied to clipboard

Challenge: Existing datasets that focus on verbal cues and focus on short-form funny videos focus on focusing on verbs and visual cue.
Approach: They curate a user-generated dataset of 10K multimodal funny videos from YouTube and annotate each video with timestamps and explanations for funny moments.
Outcome: The proposed dataset improves the ability of large language models to understand humor.
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios.
Approach: They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity.
Outcome: The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues.
Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest (2023.acl-long)

Copied to clipboard

Challenge: Large neural networks can generate jokes, but do they really “understand” humor? a new challenge challenges AI models to match a joke to a cartoon, identify a winning caption, and explain why a winner is funny.
Approach: They propose three tasks based on the New Yorker Cartoon Caption Contest . they aim to match a joke to a cartoon, identify a winning caption and explain why it's funny .
Outcome: The proposed tasks are based on the New Yorker Cartoon Caption Contest . they include matching a joke to a cartoon, identifying a winning caption, and explaining why a funny caption is funny.
GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for video comment art are constrained by their limited modalities and insufficient categories, hindering creativity in video-based comment art creation.
Approach: They propose a benchmark that integrates video and text modalities to evaluate MLLMs’ abilities to compose video Comment art.
Outcome: The proposed framework integrates video and text modalities to evaluate MLLMs’ abilities to compose video comment art.
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics (2025.findings-emnlp)

Copied to clipboard

Challenge: PixelHumor is a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs’ ability to interpret multimodal humor and recognize narrative sequences.
Approach: PixelHumor is a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs’ ability to interpret multimodal humor and recognize narrative sequences.
Outcome: Experiments with state-of-the-art LMMs reveal that top models achieve only 61% accuracy in panel sequencing, far below human performance.
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Despite advances in artificial intelligence, building social intelligence remains a challenge.
Approach: They propose a task to explain why people laugh in a video and a dataset to do this.
Outcome: The proposed dataset generates plausible explanations for laughter in video and in-the-wild videos.
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos (2025.acl-long)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) are used for video quality assessment, image captioning and video analysis.
Approach: They propose a benchmark to evaluate MLLMs on AIGC videos using coherence validation, error awareness, error type detection and reasoning evaluation tasks.
Outcome: The proposed benchmark evaluates 13 frontier MLLMs on AIGC videos.
SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to understanding laughter or humor focus on narrowly defined tasks such as detecting humor and estimating humor intensity.
Approach: They propose a dataset for real-world laughter understanding with multimodal textual representations and question–answer annotations.
Outcome: The proposed framework outperforms baselines in three laughter-related tasks, showing that it is robust.
UR-FUNNY: A Multimodal Language Dataset for Understanding Humor (D19-1)

Copied to clipboard

Challenge: Humor is a unique and creative communicative behavior often displayed during social interactions.
Approach: They present a dataset that allows to model multimodal language used in expressing humor using text, visual and acoustic communication.
Outcome: The proposed framework opens the door to understanding multimodal language used in expressing humor.
Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio (2026.acl-long)

Copied to clipboard

Challenge: Recent multimodal large language models lack robust audio-visual integration ability and performance on DeafTest is highly correlated with AV-Odyssey accuracy.
Approach: They propose a benchmarking tool that integrates audio-visual reasoning with audio-video cues to infer solutions.
Outcome: The proposed model performs well on DeafTest, but lacks audio perception in simple audio tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations