Challenge: Existing studies on empathy detection in video and audio have relied on scripted or semi-scripted interactions that fail to capture the complexities and nuances of real-life interactions.
Approach: They propose to develop a multimodal language model that detects empathy in audiovisual data by using neural architecture search and optimisation techniques.
Outcome: The proposed model will be able to detect empathy in audiovisual data and use neural architecture search to deliver it.

Similar Papers

Multimodal large language models for inclusive collaboration learning tasks (2022.naacl-srw)

Copied to clipboard

Challenge: This project leverages advances in multimodal large language models to build an inclusive collaboration feedback loop for participants developing general collaboration skills.
Approach: They propose to integrate advances in multimodal large language models into downstream tasks such as the learning analytics feedback loop.
Outcome: The proposed model will be used to detect, model, and feedback participants developing general collaboration skills.
MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model (2024.emnlp-main)

Copied to clipboard

Challenge: Existing MLLMs have a visual question answering capability but lack domain-specific information.
Approach: They propose a framework for language model modules in MLLMs when handling projected image features and verify this hypothesis using logit lens.
Outcome: The proposed framework will yield a 10% change in accuracy at most, shedding light on the development of cross-domain, all-encompassing MLLMs in the future.
MM-SOC: Benchmarking Multimodal Large Language Models in Social Media Platforms (2024.findings-acl)

Copied to clipboard

Challenge: Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emotions associated with interactions in online spaces.
Approach: They propose a benchmark to evaluate MLLMs' understanding of multimodal social media content and a large-scale YouTube tagging dataset to evaluate their performance.
Outcome: The proposed model performs better in a zero-shot setting, suggesting potential improvements.
EmpathicStories++: A Multimodal Dataset for Empathy Towards Personal Experiences (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for empathy modeling are limited in the ways they are not captured in the wild.
Approach: They propose a multimodal dataset for empathy during personal experience sharing that contains 53 hours of video, audio, and text data of 41 participants.
Outcome: The EmpathicStories++ dataset contains 53 hours of video, audio, and text data of 41 participants sharing vulnerable experiences and reading empathically resonant stories with an AI agent.
Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements (2023.findings-emnlp)

Copied to clipboard

Challenge: Empathetic dialogue is an essential part of building harmonious social relationships and contributes to the development of a helpful AI.
Approach: They propose three methods to improve the performance of large language models (LLMs) they propose semantically similar in-context learning, two-stage interactive generation and combination with the knowledge base.
Outcome: The proposed methods achieve state-of-the-art in automatic and human evaluations and the possibility of GPT-4 simulating human evaluators.
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Empathy plays a pivotal role in fostering prosocial behavior, often triggered by the sharing of personal experiences through narratives.
Approach: They propose to use contrastive learning with masked LMs and supervised fine-tuning with large language models to improve empathy understanding in NLP models.
Outcome: The proposed methods show that there is low agreement among annotators and that cultural differences are a factor in their interpretation of empathy.
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances extend language understanding beyond text to speech, enabling unified reasoning across modalities.
Approach: They construct and release a speech-augmented benchmark based on Global MMLU Lite and a data set spanning English, Chinese, and Korean.
Outcome: The proposed model is robust to demographic factors but sensitive to language and option order, suggesting that speech can amplify structural biases.
Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection (2025.emnlp-main)

Copied to clipboard

Challenge: a new study examines the effectiveness of large language models and non-LLMs in multimodal intent detection . large-scale multimodal data integrations include text, audio, and visual inputs .
Approach: They propose a framework to debias multimodal intent detection datasets by using human evaluation.
Outcome: The proposed framework debiases the datasets and shows that mistral-7B outperforms most competitive models by approximately 9% on MIntRec-1 and 4% on MIndRec2.0.
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives (2024.findings-acl)

Copied to clipboard

Challenge: Existing video-language understanding systems with human-like senses can mimic both our linguistic medium and visual environment with temporal dynamics.
Approach: They propose to develop video-language understanding systems with human-like senses . they summarize their methods and highlight challenges associated with them .
Outcome: The proposed models perform well in a variety of tasks and domains.
The Pursuit of Empathy: Evaluating Small Language Models for PTSD Dialogue Support (2025.emnlp-main)

Copied to clipboard

Challenge: Claude Sonnet 3.5 consistently outperforms all models, but smaller models often approach human-rated empathy levels.
Approach: They introduce a dataset comprising 10,000 two-turn conversations across 500 diverse, clinically-grounded PTSD personas.
Outcome: The proposed model outperforms all models but has a "knowledge transfer ceiling" older adults prefer validation responses while graduate-educated users prefer emotionally layered responses .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations