Challenge: Existing multimodal dialogue systems are based on unimodal sources, capturing information from text and image.
Approach: They propose a position and attribute aware attention mechanism to learn enhanced image representation conditioned on the user utterance.
Outcome: The proposed model outperforms the state-of-the-art models on text similarity metrics.

Similar Papers

From Traits to Empathy: Personality-Aware Multimodal Empathetic Response Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing approaches focus on acquiring affective and cognitive knowledge from text, but neglect the unique personality traits of individuals and the inherently multimodal nature of human face-to-face conversation.
Approach: They propose a multimodal dialogue system that generates empathetic responses from a perspective that considers the personality traits of users.
Outcome: The proposed system generates empathetic responses from a multimodal perspective and analyzes multimodal data to understand the user’s emotional state and situation.
Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for dialogue-to-image retrieval are constrained by pre-trained vision language models.
Approach: They leverage the reasoning capabilities of large language models to predict potential features in images to be shared based on dialogue context.
Outcome: The proposed method outperforms existing methods significantly in terms of Recall@k.
Multimodal Dialogue Response Generation (2022.acl-long)

Copied to clipboard

Challenge: Existing studies focus on multimodal dialogue models but neglect generation methods.
Approach: They propose a multimodal dialogue response generation task which requires multimodal dialogs containing both texts and images which are difficult to obtain.
Outcome: Experiments show that the proposed model can generate informative text and high-resolution image responses.
Large Scale Generative Multimodal Attribute Extraction for E-commerce Attributes (2023.acl-industry)

Copied to clipboard

Challenge: E-commerce websites often don’t label or mislabel attributes of products .
Approach: They propose a multi-modal product attribute generation system that extracts product attributes from the product pages of eCommerce stores by using both text and images.
Outcome: The proposed model improves the recall@90P accuracy by 10.16% and 6.9 from the state-of-the-art models.
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities.
Approach: They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs.
Outcome: The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content.
Data Augmentation with Paraphrase Generation and Entity Extraction for Multimodal Dialogue System (2022.lrec-1)

Copied to clipboard

Challenge: Contextually aware intelligent agents are often required to understand the users and their surroundings in real-time.
Approach: They propose to build a multimodal dialogue system for children learning basic math concepts using limited datasets.
Outcome: The proposed system improves the Natural Language Understanding (NLU) module of a task-oriented SDS pipeline with limited dataset resources.
Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems (2022.acl-long)

Copied to clipboard

Challenge: Existing work on empathetic dialogues focused on the two-party scenario, but multi-party dialogues are pervasive in reality.
Approach: They propose a multi-party empathetic dialogue generation task that uses a static-dynamic model to explore emotion and sensibility.
Outcome: The proposed task is based on a model with static sensibility and dynamic emotion . it achieves state-of-the-art performance in multi-party empathetic dialogue learning .
MultiDM-GCN: Aspect-guided Response Generation in Multi-domain Multi-modal Dialogue System using Graph Convolutional Network (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing research suggests that engaging conversations include visual cues (e.g., a video or images) or audio cue.
Approach: They propose a multi-modal conversational framework that generates the responses following the different aspects of a product or service to cater to the user's needs.
Outcome: The proposed framework outperforms baselines for the task-oriented dialogue setup.
Multi-Modal Open-Domain Dialogue (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work in open-domain conversational agents has demonstrated that significant improvements in humanness and user preference can be achieved via massive scaling in both pre-training data and model size.
Approach: They combine open-domain dialogue agents with vision models to investigate human preferences and humanness.
Outcome: The proposed model outperforms existing models in multi-modal dialogue while performing as well as its predecessor (text-only) BlenderBot.
TIGER: A Unified Generative Model Framework for Multimodal Dialogue Response Generation (2024.lrec-main)

Copied to clipboard

Challenge: Existing research on multimodal dialogues focuses on textual response generation and visual response selection based on the dialogue context.
Approach: They propose a generative model framework for multimodal dialogue response generation that ground the conversation on an image.
Outcome: The proposed system provides users with an enhanced conversational experience.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations