Challenge: Existing studies focus on image-sharing behavior in singular sessions, leading to limited long-term social interaction.
Approach: They propose a large-scale long-term multi-modal dialogue dataset that generates long-time multi-modity dialogue distilled from ChatGPT and proposed image aligner.
Outcome: The proposed framework generates long-term multi-modal dialogue from ChatGPT and image aligner.

Similar Papers

SHARE: Shared Memory-Aware Open-Domain Long-Term Dialogue Dataset Constructed from Movie Script (2025.acl-long)

Copied to clipboard

Challenge: Antoine de Saint-Exupéry Memory in dialogue plays a crucial role in building relationships and facilitating the ongoing conversation.
Approach: They propose a long-term dialogue dataset named SHARE that includes shared memories between two individuals.
Outcome: The proposed dataset makes long-term dialogues more engaging and sustainable . it includes summaries of persona information and events of two individuals .
MMChat: Multi-Modal Chat Dataset on Social Media (2022.lrec-1)

Copied to clipboard

Challenge: Incorporating multi-modal contexts in conversation is important for developing engaging dialogue systems.
Approach: They propose a large scale Chinese multi-modal dialogue corpus that contains image-grounded dialogues from real conversations on social media.
Outcome: The proposed model can handle sparsity issues in dialogue generation tasks by incorporating image features.
Multi-Modal Open-Domain Dialogue (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work in open-domain conversational agents has demonstrated that significant improvements in humanness and user preference can be achieved via massive scaling in both pre-training data and model size.
Approach: They combine open-domain dialogue agents with vision models to investigate human preferences and humanness.
Outcome: The proposed model outperforms existing models in multi-modal dialogue while performing as well as its predecessor (text-only) BlenderBot.
Commonsense-augmented Memory Construction and Management in Long-term Conversations via Context-aware Persona Refinement (2024.eacl-short)

Copied to clipboard

Challenge: Memorizing and utilizing speakers’ personas is a common practice for response generation in long-term conversations, yet human-authored datasets often provide uninformative persona sentences that hinder response quality.
Approach: They propose a framework that leverages commonsense-based persona expansion to address such issues in long-term conversations.
Outcome: The proposed framework facilitates better response generation via human-like persona refinement.
Constructing Multi-Modal Dialogue Dataset by Replacing Text with Semantically Relevant Images (2021.acl-short)

Copied to clipboard

Challenge: Existing training methods for multi-modal dialogue systems rely on image captioning or visual question answering datasets that are irrelevant to the dialogue context.
Approach: They propose to create a 45k multi-modal dialogue dataset with minimal human intervention . they use text dialogue datasets, image-mixed dialogues and contextual-similarity filtering .
Outcome: The proposed dataset can be used as training data for multi-modal dialogue systems . human evaluations show that the model can be effectively used .
Post Persona Alignment for Multi-Session Dialogue Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for multi-session persona-based dialogue generation typically retrieve persona information before response generation, which can constrain diversity and result in generic outputs.
Approach: They propose a two-stage framework that reverses the process of retrieving persona information before response generation.
Outcome: Experiments on multi-session persona-based dialogue data show that the proposed framework outperforms existing methods in consistency, diversity, and persona relevance.
Long Time No See! Open-Domain Conversation with Long-Term Persona Memory (2022.findings-acl)

Copied to clipboard

Challenge: Existing persona dialogue datasets and models can build long-term relationships with humans . however, current open-domain dialogue systems cannot build long relationships with users .
Approach: They propose a long-term memory conversation dataset and a dialogue generation framework with long-Term memory mechanism to extract and continuously update long-time persona memory.
Outcome: The proposed system outperforms baselines in terms of long-term dialogue consistency . the proposed system can build long-lasting relationships between humans and bots .
LongMP-Bench: A Benchmark for Multimodal Persona Understanding in Long-Term Dialogues (2026.findings-acl)

Copied to clipboard

Challenge: Existing datasets suffer from limited persona diversity and static, overly simplified settings, making them insufficient for capturing the complexity of real-world interactions.
Approach: They propose a benchmark to evaluate models' ability to understand evolving user personas within long-term multimodal dialogues by using a dataset that contains long conversations from 150 users.
Outcome: The proposed benchmark aims to assess models' ability to track persona evolution, integrate visual and textual inputs, and apply persona understanding in realistic dialogue scenarios.
MPCHAT: Towards Multimodal Persona-Grounded Conversation (2023.acl-long)

Copied to clipboard

Challenge: Existing research on persona-based dialogue has focused on textual persona that delivers personal facts or personalities, but image modality can reveal the speaker’s personal characteristics and experiences in episodic memory.
Approach: They propose a multimodal persona-based dialogue dataset which extends persona with both text and images to contain episodic memories.
Outcome: The proposed dataset extends persona with text and images to contain episodic memories.
DialogueTRM: Exploring Multi-Modal Emotional Dynamics in a Conversation (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on the self and inter-personal dependencies in multi-modal conversations, but they ignore the temporal and spatial dependencies.
Approach: They propose a Dialogue Transformer for simultaneously modeling the intra-modal and inter-modal emotion dynamics.
Outcome: The proposed models outperform the state-of-the-art on three benchmark datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations