MMChat: Multi-Modal Chat Dataset on Social Media (2022.lrec-1)

Copied to clipboard

Challenge: Incorporating multi-modal contexts in conversation is important for developing engaging dialogue systems.
Approach: They propose a large scale Chinese multi-modal dialogue corpus that contains image-grounded dialogues from real conversations on social media.
Outcome: The proposed model can handle sparsity issues in dialogue generation tasks by incorporating image features.

Similar Papers

MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain Conversation (2023.acl-long)

Copied to clipboard

Challenge: MMDialog is a dataset of 1.08 million real-world dialogues with 1.53 million unique images across 4,184 topics.
Approach: They propose to use a curated set of 1.08 million dialogues with 1.53 million unique images to generalize the open domain.
Outcome: The proposed system can predict responses to multi-modal content with state-of-the-art techniques and measure their performance.
Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on image-sharing behavior in singular sessions, leading to limited long-term social interaction.
Approach: They propose a large-scale long-term multi-modal dialogue dataset that generates long-time multi-modity dialogue distilled from ChatGPT and proposed image aligner.
Outcome: The proposed framework generates long-term multi-modal dialogue from ChatGPT and image aligner.
Constructing Multi-Modal Dialogue Dataset by Replacing Text with Semantically Relevant Images (2021.acl-short)

Copied to clipboard

Challenge: Existing training methods for multi-modal dialogue systems rely on image captioning or visual question answering datasets that are irrelevant to the dialogue context.
Approach: They propose to create a 45k multi-modal dialogue dataset with minimal human intervention . they use text dialogue datasets, image-mixed dialogues and contextual-similarity filtering .
Outcome: The proposed dataset can be used as training data for multi-modal dialogue systems . human evaluations show that the model can be effectively used .
MPCHAT: Towards Multimodal Persona-Grounded Conversation (2023.acl-long)

Copied to clipboard

Challenge: Existing research on persona-based dialogue has focused on textual persona that delivers personal facts or personalities, but image modality can reveal the speaker’s personal characteristics and experiences in episodic memory.
Approach: They propose a multimodal persona-based dialogue dataset which extends persona with both text and images to contain episodic memories.
Outcome: The proposed dataset extends persona with text and images to contain episodic memories.
MultiDM-GCN: Aspect-guided Response Generation in Multi-domain Multi-modal Dialogue System using Graph Convolutional Network (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing research suggests that engaging conversations include visual cues (e.g., a video or images) or audio cue.
Approach: They propose a multi-modal conversational framework that generates the responses following the different aspects of a product or service to cater to the user's needs.
Outcome: The proposed framework outperforms baselines for the task-oriented dialogue setup.
LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming (2023.acl-long)

Copied to clipboard

Challenge: a recent study shows that open-domain dialogue systems are not able to perform well in fast-growing scenarios such as live streaming due to the domain gap between online-post constructed data and those required in downstream conversational tasks.
Approach: They propose to train a conversational agent based on large social media datasets with multiple domains to improve response in live streaming scenarios.
Outcome: The proposed model improves response modeling and addressee recognition in live open-domain scenarios.
MMCoQA: Conversational Question Answering over Text, Tables, and Images (2022.acl-long)

Copied to clipboard

Challenge: Existing conversational QA systems only use a single knowledge source, e.g., paragraphs or knowledge graph, and assume it contains enough evidence to extract answers to users' questions.
Approach: They propose a task to answer users' questions with multimodal knowledge sources via multi-turn conversations using a multimodal dataset.
Outcome: The proposed task brings a series of research challenges, including but not limited to priority, consistency, and complementarity of multimodal knowledge.
MAGID: An Automated Pipeline for Generating Synthetic Multi-modal Datasets (2024.naacl-long)

Copied to clipboard

Challenge: Existing approaches to augment textual dialogues with retrieved images pose privacy, diversity, and quality constraints.
Approach: They propose a framework to augment text-only dialogues with diverse and high-quality images by using a diffusion model and a feedback loop.
Outcome: The proposed framework is comparable to or better than baselines, with significant improvements in human evaluation, especially against retrieval baselines where the image database is small.
M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue Database (2022.acl-long)

Copied to clipboard

Challenge: Existing data resources to support multimodal affective analysis in dialogues are limited in scale and diversity.
Approach: They propose a multimodal multi-scene multi-label Emotional Dialogue dataset, M3ED, which contains 990 dyadic emotional dialogues from 56 different TV series.
Outcome: The proposed dataset contains 990 dyadic emotional dialogues from 56 different TV series, a total of 9,082 turns and 24,449 utterances.
Game-Based Video-Context Dialogue (D18-1)

Copied to clipboard

Challenge: Current dialogue systems focus more on textual and speech context knowledge and are usually based on two speakers.
Approach: They propose to use live soccer game videos and Twitch.tv chats to develop visual-grounded dialogue models.
Outcome: The proposed model can generate relevant temporal and spatial event language from live video and chat history while also being relevant to chat history.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations