A Large-Scale Corpus for Conversation Disentanglement (P19-1)

Copied to clipboard

Challenge: a dataset of 77,563 messages manually annotated with reply-structure graphs disentangles conversations and defines internal conversation structure.
Approach: They use a dataset of 77,563 messages manually annotated with reply-structure graphs to disentangle conversations and define internal conversation structure.
Outcome: The new dataset is 16 times larger than all previous datasets combined and includes adjudication of annotation disagreements and context.

Similar Papers

Structural Characterization for Dialogue Disentanglement (2022.acl-long)

Copied to clipboard

Challenge: tangled multi-party dialogues lead to difficulties in understanding the dialogue history for both human and machine.
Approach: They propose a model for disentangling multi-party dialogues using speaker property and reference dependency.
Outcome: The proposed model achieves state-of-the-art on the Ubuntu IRC benchmark dataset and contributes to dialogue-related comprehension.
Conversation Disentanglement with Bi-Level Contrastive Learning (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focus on pairwise utterance relations but pay inadequate attention to utterant-to-context relation modeling.
Approach: They propose a general disentangle model based on bi-level contrastive learning that brings closer utterances in the same session while encouraging each utterrance to be near its clustered session prototypes in representation space.
Outcome: The proposed model achieves state-of-the-art performance on both settings across public datasets.
A Large-Scale Corpus of E-mail Conversations with Standard and Two-Level Dialogue Act Annotations (2020.coling-main)

Copied to clipboard

Challenge: e-mail conversations have domain-agnostic and two-level dialogue act annotations . et al. (2017): a better understanding of asynchronous conversations.
Approach: They present a large-scale corpus of e-mail conversations with domain-agnostic and two-level dialogue act annotations . they use ISO standard 24617-2 as the annotation scheme to annotate over 6,000 messages and 35,000 sentences .
Outcome: The proposed model outperforms other neural networks but falls short of human performance.
AUGUST: an Automatic Generation Understudy for Synthesizing Conversational Recommendation Datasets (2023.findings-acl)

Copied to clipboard

Challenge: Existing work on conversational recommendation systems lacks high-quality data . existing datasets lack large-scale and high-level data based on human annotators .
Approach: They propose an automatic dataset synthesis approach that generates large-scale recommendation dialogues using structured graphs based on user-item information from the real world.
Outcome: The proposed approach can generate large-scale and high-quality recommendation dialogues . it exploits user preferences, knowledge graphs, and conversation ability from existing datasets based on real-world data .
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information.
Approach: They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants.
Outcome: The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants.
Unsupervised Conversation Disentanglement through Co-Training (2021.emnlp-main)

Copied to clipboard

Challenge: Existing work on conversation disentanglement relies heavily on human annotations, which is expensive to obtain in practice.
Approach: They propose to train a conversation disentanglement model without referencing human annotations . they use a message-pair classifier and a session classifier to retrieve local relations .
Outcome: The proposed method achieves competitive performance compared to previous methods on a large movie dialogue dataset.
Summarization of Dialogues and Conversations At Scale (2023.eacl-tutorials)

Copied to clipboard

Challenge: Conversations are the natural communication format for people.
Approach: This tutorial will survey the cutting-edge methods for summarizing written and spoken conversation.
Outcome: This tutorial will examine the cutting-edge methods for summarizing written and spoken conversations, covering key sub-areas whose combination is needed for a successful solution.
SuperDialseg: A Large-scale Dataset for Supervised Dialogue Segmentation (2023.emnlp-main)

Copied to clipboard

Challenge: Empirical studies show that supervised learning is extremely effective in in-domain datasets and models trained on SuperDialseg can achieve good generalization ability on out-of-domain data.
Approach: They propose a supervised definition of dialogue segmentation points using document-grounded dialogues and a large-scale supervised dataset called SuperDialseg.
Outcome: The proposed model can achieve good generalization ability on out-of-domain data.
More Diverse Dialogue Datasets via Diversity-Informed Data Collection (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to generate conversational dialogue produce uninteresting, predictable responses.
Approach: They propose a method to collect and determine more diverse data from conversational participants . they use dynamically computed corpus-level statistics to determine which conversational participant to collect data from .
Outcome: The proposed method produces significantly more diverse data than baseline methods and better results on emotion classification and dialogue generation tasks.
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization (D19-54)

Copied to clipboard

Challenge: Existing work on abstractive dialogue summarizations has focused on news summarizing but there is no such comprehensive dataset.
Approach: They propose to use a chat-dialogues corpus with abstractive dialogue summaries to generate a short version of text that covers the main points succinctly.
Outcome: The proposed dataset achieves higher ROUGE scores than the model-generated summaries of news, compared with human evaluators' judgement.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations