Challenge: Currently, video-based dialogue systems rely on a single dialogue type, hindering their versatility in practical applications.
Approach: They propose to generate video-driven multilingual mixed-type dialogues using KwaiChat . they propose to create a video-based multilingual mix of 4 dialogue types, 30 domains, 4 languages, 13 topics .
Outcome: The proposed model performs best on KwaiChat but is not perfect in this situation.

Similar Papers

Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models (2024.acl-long)

Copied to clipboard

Challenge: a surge of deep learning applications for video understanding have led to major advancements in video-related tasks.
Approach: They propose a multimodal video-based conversation model that merges a video-adapted visual encoder with an LLM and a dataset that is easily scalable and robust to label noise.
Outcome: The proposed model can understand and generate detailed conversations about videos.
MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling (D18-1)

Copied to clipboard

Challenge: a dataset of 10k human-human written conversations is one order of magnitude larger than previous annotated task-oriented corpora.
Approach: They propose to collect 10k human-human written conversations from a crowd-sourced dataset using crowd-sourcing.
Outcome: The proposed dataset is one order of magnitude larger than previous annotated task-oriented corpora and shows the usability of the data and sets a baseline for future studies.
LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming (2023.acl-long)

Copied to clipboard

Challenge: a recent study shows that open-domain dialogue systems are not able to perform well in fast-growing scenarios such as live streaming due to the domain gap between online-post constructed data and those required in downstream conversational tasks.
Approach: They propose to train a conversational agent based on large social media datasets with multiple domains to improve response in live streaming scenarios.
Outcome: The proposed model improves response modeling and addressee recognition in live open-domain scenarios.
Game-Based Video-Context Dialogue (D18-1)

Copied to clipboard

Challenge: Current dialogue systems focus more on textual and speech context knowledge and are usually based on two speakers.
Approach: They propose to use live soccer game videos and Twitch.tv chats to develop visual-grounded dialogue models.
Outcome: The proposed model can generate relevant temporal and spatial event language from live video and chat history while also being relevant to chat history.
Video-Grounded Dialogues with Pretrained Generation Language Models (2020.acl-main)

Copied to clipboard

Challenge: Pre-trained language models have shown success in improving downstream NLP tasks . pre-tuned models capture textual dependencies in text data of rich semantics .
Approach: They propose a framework for improving video-grounded dialogue by extending GPT-2 models . they propose to combine visual and textual representation into a structured sequence .
Outcome: The proposed framework improves audio-visual scene-aware dialogues benchmark on AVSD . it is based on a large pre-trained GPT-2 network and can generate natural responses .
XDailyDialog: A Multilingual Parallel Dialogue Corpus (2023.acl-long)

Copied to clipboard

Challenge: Existing datasets for open-domain dialogue modeling limited to a single language . absence of multilingual datasets hinders development of robust open- domain dialog systems .
Approach: They propose a multilingual parallel open-domain dialog dataset to explore multilingual and cross-lingual open- domain dialog.
Outcome: The proposed model can be used to explore multilingual and cross-lingual open-domain dialogs in other languages.
CrossWOZ: A Large-Scale Chinese Cross-Domain Task-Oriented Dialogue Dataset (2020.tacl-1)

Copied to clipboard

Challenge: Despite the significant contributions to the community, there is still a gap between existing dialogue corpora and real-life human dialogue data.
Approach: They propose to develop Chinese cross-domain wizard-of-oz task-oriented dataset CrossWOZ with rich annotations of dialogue states and dialogue acts on both user and system sides.
Outcome: The proposed dataset contains 6K dialogue sessions and 102K utterances for 5 domains, including hotel, restaurant, attraction, metro, and taxi.
ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have led to significant breakthroughs in the field.
Approach: They evaluate ChatGPT over multiple tasks with diverse languages and large datasets to provide more comprehensive information for multilingual NLP applications.
Outcome: The proposed model can process and generate texts for multiple languages due to its multilingual training data.
Automatic Generation of Large-scale Multi-turn Dialogues from Reddit (2022.coling-1)

Copied to clipboard

Challenge: Using a set of algorithms, we can generate large dialogue corpus from Reddit.
Approach: They propose to automatically convert posts and their comments from discussion forums such as Reddit into multi-turn dialogues.
Outcome: The proposed methods improve on the baseline method by 36.3% . the best method shows an improvement of 36.6% over the previous one .
Natural Language Processing for Multilingual Task-Oriented Dialogue (2022.acl-tutorials)

Copied to clipboard

Challenge: a tutorial will examine the challenges and gaps in multilingual ToD research . multilingual systems are difficult to build, and are limited to English and other languages .
Approach: This tutorial will discuss the importance of multilingual task-oriented dialogue systems . it will provide an overview of current research gaps, challenges and initiatives related to multilingual ToD systems - with a particular focus on their connections to current research and challenges in multilingual and low-resource NLP.
Outcome: This tutorial will provide an overview of current research gaps, challenges and initiatives related to multilingual ToD systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations