Challenge: Existing benchmarks fail to achieve ecological validity, signal clarity, and reliable fine-grained labeling in multimodal Emotion Recognition (MER) Existing datasets lack spontaneity of real-life interactions, resulting in poor quality and inconsistent data quality.
Approach: They propose a bilingual benchmark to resolve limitations of ecological validity and noise in existing datasets by combining strictly filtered static slices with a dynamic Streaming Monologue subset.
Outcome: EmoS provides trusted ground truth that captures continuous emotional evolution.

Similar Papers

EmoBench: Evaluating the Emotional Intelligence of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for Emotional Intelligence (EI) focus on emotion recognition, neglecting essential EI capabilities.
Approach: They propose a benchmark that proposes a comprehensive definition for machine EI . they propose 400 hand-crafted questions in English and Chinese to evaluate EI.
Outcome: The proposed benchmarks focus on emotion recognition, neglecting EI capabilities . they are constructed from existing datasets, which include frequent patterns and errors . the proposed benchmark includes questions in English and Chinese that require thorough reasoning and understanding .
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description (2025.findings-naacl)

Copied to clipboard

Challenge: Existing 3D facial emotion modeling models are constrained by limited emotion classes and insufficient datasets.
Approach: They propose a 3D facial emotion modeling dataset that spans a wide spectrum of human emotions . they use large language models to generate a diverse array of textual descriptions .
Outcome: Emo3D is an extensive dataset that spans human emotions with images and 3D blendshapes.
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation (2025.naacl-long)

Copied to clipboard

Challenge: Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, but many benchmarks suffer from systematic biases.
Approach: They propose a benchmark to avoid Type-I errors by creating one perception question and one knowledge anchor question through a meticulous annotation process.
Outcome: The proposed benchmark avoids Type-I errors while maintaining reliability of MCQ evaluations.
Emo Pillars: Knowledge Distillation to Support Fine-Grained Context-Aware and Context-Less Emotion Classification (2025.findings-acl)

Copied to clipboard

Challenge: a recent study shows that sentiment analysis datasets lack context in which an opinion was expressed and are limited by a few emotion categories.
Approach: They propose to ground an LLM-based model into a corpus of narratives to generate stories-character-centered utterances with unique contexts over 28 emotion classes.
Outcome: The proposed model generates non-repetitive story-character-centered utterances with unique contexts over 28 emotion classes.
AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models are increasingly used in Personalized Image Aesthetic Assessment (PIAA) however, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education.
Approach: They propose to evaluate MLLMs along two complementary dimensions: (1) stereotype bias and (2) alignment between model outputs and genuine human aesthetic preferences.
Outcome: The proposed benchmark covers three subtasks: aesthetic perception, assessment, empathy and alignment between outputs and genuine human aesthetic preferences.
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations (2026.findings-acl)

Copied to clipboard

Challenge: Existing datasets face issues such as low quality, limited scale, and incomplete modalities, hindering model performance.
Approach: They propose to use Chinese multimodal datasets to capture authentic emotional interplay from 19 professional actors.
Outcome: The EmotionTalk dataset spans 23.6 hours of dyadic conversations across diverse scenarios.
EmoRes: Toward Adaptive Psychological Support via User-Agnostic Benchmark and Topic-Mining Agent (2026.findings-acl)

Copied to clipboard

Challenge: Large language models generate fragmented and emotionally inconsistent dialogues lacking the therapeutic structure necessary for reliable assessment.
Approach: They propose a framework that boosts psychological reasoning via a Topic-Mining Emotional Agent and a multi-perspective Self-Reflection Agent.
Outcome: The proposed framework improves topic continuity, emotional coherence, and clinical interpretability over baselines and validated by ablation studies and human evaluations.
EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues (2025.naacl-long)

Copied to clipboard

Challenge: EmoCharacter evaluates emotional fidelity of role-playing agents in dialogues . current evaluations focus on personality fidelity, tone imitation, and knowledge consistency .
Approach: They propose a benchmark to assess emotional fidelity of role-playing agents in dialogues using large language models.
Outcome: The proposed benchmark measures emotional fidelity of role-playing agents and the characters they portray.
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships? (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on assessing factual and logical correctness in downstream tasks with limited emphasis on evaluating MLLMs’ ability to interpret pragmatic cues and intermodal relationships.
Approach: They propose to use Coherence Relations to assess MLLMs' ability to perform multimodal discourse analysis using different prompting strategies.
Outcome: The proposed model fails to match the performance of simple classifier-based benchmarks on 10+ MLLMs using different prompting strategies.
Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses (2026.acl-long)

Copied to clipboard

Challenge: Culture is a fundamental determinant of human affective processing and affective perceptions are often limited by declarative knowledge or established societal customs.
Approach: They propose a multimodal benchmark that leverages LLM-generated provisional labels to isolate cross-cultural emotional distinctions.
Outcome: The proposed benchmark captures cross-cultural emotional distinctions and derives reliable ground-truth annotations through human evaluation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations