Challenge: Multimodality has been explored in multi-party and multi-session conversations, but task-specific constraints have hindered its seamless integration into dynamic, natural conversations.
Approach: They propose a multimodal conversation dataset and a model with multimodal memory retrieval to equip chatbots with "eyes and ears" they aim to integrate multimodality into chatbot interactions by integrating visual and auditory inputs into the chatbot.
Outcome: The proposed model demonstrates the ability to engage in long-term conversations with multiple speakers in complex, real-world-like settings, effectively processing visual and auditory inputs to understand and respond appropriately.

Similar Papers

Multi-party Multimodal Conversations Between Patients, Their Companions, and a Social Robot in a Hospital Memory Clinic (2024.eacl-demo)

Copied to clipboard

Challenge: a new spoken dialogue system is being developed for hospitals and hospitals to enable multi-party interactions . a social robot can be used to have multi-part conversations with patients and their companions .
Approach: They describe a spoken dialogue system that allows patients to have multi-party conversations with their companions . they use speech and video input to generate both speech and gestures - arm, head, and eye movements .
Outcome: The proposed system generates human-like clarification requests when the patient pauses mid-utterance, answers in-domain questions, and responds appropriately to out-of-domain requests.
Multi-Modal Open-Domain Dialogue (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work in open-domain conversational agents has demonstrated that significant improvements in humanness and user preference can be achieved via massive scaling in both pre-training data and model size.
Approach: They combine open-domain dialogue agents with vision models to investigate human preferences and humanness.
Outcome: The proposed model outperforms existing models in multi-modal dialogue while performing as well as its predecessor (text-only) BlenderBot.
ICON: Interactive Conversational Memory Network for Multimodal Emotion Detection (D18-1)

Copied to clipboard

Challenge: Existing studies do not explicitly consider inter-personal influences that thrive in the emotional dynamics of dialogues.
Approach: They propose a multimodal emotion detection framework that extracts multimodal features from conversational videos and hierarchically models the self- and inter-speaker emotional influences into global memories.
Outcome: The proposed model outperforms state-of-the-art networks on multiple classification and regression tasks in two benchmark datasets.
Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal Emotion Recognition in Conversations (MERC) is a new way to enhance human-computer interaction.
Approach: This survey offers a systematic overview of Multimodal Emotion Recognition in Conversations . it examines motivations, core tasks, representative methods, and evaluation strategies .
Outcome: The survey examines the effectiveness of MERC and its evaluation strategies.
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling (2025.findings-acl)

Copied to clipboard

Challenge: Conversational assistants are increasingly popular across diverse real-world applications . speech data constitute high-dimensional signals that are difficult to model even for frontier models .
Approach: They propose a data-centric customization approach for enhancing multimodal understanding in conversational speech modeling.
Outcome: The proposed model achieves state-of-the-art on the Spoken-SQuAD benchmark using 10% of training data with open-weight models.
MMCoQA: Conversational Question Answering over Text, Tables, and Images (2022.acl-long)

Copied to clipboard

Challenge: Existing conversational QA systems only use a single knowledge source, e.g., paragraphs or knowledge graph, and assume it contains enough evidence to extract answers to users' questions.
Approach: They propose a task to answer users' questions with multimodal knowledge sources via multi-turn conversations using a multimodal dataset.
Outcome: The proposed task brings a series of research challenges, including but not limited to priority, consistency, and complementarity of multimodal knowledge.
MPCHAT: Towards Multimodal Persona-Grounded Conversation (2023.acl-long)

Copied to clipboard

Challenge: Existing research on persona-based dialogue has focused on textual persona that delivers personal facts or personalities, but image modality can reveal the speaker’s personal characteristics and experiences in episodic memory.
Approach: They propose a multimodal persona-based dialogue dataset which extends persona with both text and images to contain episodic memories.
Outcome: The proposed dataset extends persona with text and images to contain episodic memories.
DialogueTRM: Exploring Multi-Modal Emotional Dynamics in a Conversation (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on the self and inter-personal dependencies in multi-modal conversations, but they ignore the temporal and spatial dependencies.
Approach: They propose a Dialogue Transformer for simultaneously modeling the intra-modal and inter-modal emotion dynamics.
Outcome: The proposed models outperform the state-of-the-art on three benchmark datasets.
Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations (2023.emnlp-main)

Copied to clipboard

Challenge: open-domain chatbots focus on short single-session dialogue, neglecting the potential need for understanding contextual information in multiple consecutive sessions.
Approach: They propose a 1M multi-session dialogue dataset for integrating time intervals and speaker relationships into a long-term conversation setup.
Outcome: The proposed model can generate coherent responses according to time intervals and speaker relationships with high user engagement without contradiction in a long-term conversation setup.
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in vision-language-action models prioritize robotic action mastery . however, models trained on visual-text pairs struggle to interpret multimodal data .
Approach: They propose a framework that integrates multimodal data after initial control mastery and a Mixture-of-Experts architecture to minimize task interference.
Outcome: The proposed framework surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks and achieves six times higher performance on visual question-answering datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations