Challenge: Large-scale conventional dialogue corpora are mainly built for specified tasks with specially designed dialogue states.
Approach: They propose to annotate large-scale dialogue data with an extended ISO-24617-2 dialogue act tag-set to model a natural conversation with machines.
Outcome: The proposed corpus covers a wider range of dialogue tasks than existing task-oriented systems or text-chat systems.

Similar Papers

Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)

Copied to clipboard

Challenge: a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations .
Approach: They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner.
Outcome: The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings.
A Unifying View On Task-oriented Dialogue Annotation (2022.lrec-1)

Copied to clipboard

Challenge: Recent research attention in task-oriented dialogue systems focuses on end-to-end neural models.
Approach: They present a dataset that combines annotated corpora from four domains to provide a unified ontology and annotation schema for task-oriented dialogues.
Outcome: The proposed dataset improves language, information content and performance in dialogues with two recent models.
The ISO Standard for Dialogue Act Annotation, Second Edition (2020.lrec-1)

Copied to clipboard

Challenge: ISO standard 24617-2 for dialogue act annotation has been used in corpus annotation and in the design of components for spoken and multimodal interactive systems.
Approach: ISO standard 24617-2 for dialogue act annotation is proposed for a second edition . this second edition allows a more accurate annotation of dependence relations and rhetorical relations in dialogue.
Outcome: The proposed second edition of ISO 24617-2 for dialogue act annotation addresses some inaccuracies and undesirable limitations.
Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems (2021.naacl-main)

Copied to clipboard

Challenge: Existing goal-oriented dialogue datasets focus on identifying slots and values, but in reality, customer service agents follow multi-step procedures derived from explicit company policies.
Approach: They propose to use a fully-labeled dataset to study customer service dialogue systems in real-world scenarios.
Outcome: The proposed dataset outperforms existing models but still lacks 50.8% absolute accuracy to reach human-level performance on the dataset.
A Large-Scale Corpus of E-mail Conversations with Standard and Two-Level Dialogue Act Annotations (2020.coling-main)

Copied to clipboard

Challenge: e-mail conversations have domain-agnostic and two-level dialogue act annotations . et al. (2017): a better understanding of asynchronous conversations.
Approach: They present a large-scale corpus of e-mail conversations with domain-agnostic and two-level dialogue act annotations . they use ISO standard 24617-2 as the annotation scheme to annotate over 6,000 messages and 35,000 sentences .
Outcome: The proposed model outperforms other neural networks but falls short of human performance.
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information.
Approach: They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants.
Outcome: The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants.
Multilingual Dialogue Generation and Localization with Dialogue Act Scripting (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to training or evaluating non-English dialogue datasets often introduce artifacts that reduce their naturalness and cultural appropriateness.
Approach: They propose a structured framework for encoding, localizing, and generating multilingual dialogues from abstract intent representations.
Outcome: The proposed framework outperforms translation models in Italian, German, and Chinese on cultural relevance, coherence, and situational appropriateness.
The ADELE Corpus of Dyadic Social Text Conversations:Dialog Act Annotation with ISO 24617-2 (L18-1)

Copied to clipboard

Challenge: Recent studies have focused on task-based or instrumental dialogs, but there is increasing interest in social or interactional dialogs.
Approach: They describe a corpus of 193 dyadic text dialogs based on a novel 'getting to know you' social dialog elicitation paradigm and propose additional acts to better cover greeting and leavetaking.
Outcome: The proposed actions cover greeting and leavetaking, and the proposed acts improve the interaction between the dialogs and spoken language.
ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational Agents (C18-1)

Copied to clipboard

Challenge: Existing methods for DA annotation are incompatible with each other and do not cover all aspects necessary for open-domain human-machine interaction.
Approach: They propose to map publicly available corpora to a subset of the ISO standard and create a task-independent training corpus for DA classification.
Outcome: The proposed method can train a domain-independent DA tagger on out-of-domain conversational data and achieve robustness across different DA categories.
EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems (2022.lrec-1)

Copied to clipboard

Challenge: Existing emotion-annotated task-oriented corpora are limited in size, label richness, and public availability, creating a bottleneck for downstream tasks.
Approach: They propose a large-scale manually emotion-annotated corpus of task-oriented dialogues based on a multi-domain task-orientated dataset.
Outcome: The proposed method is based on a task-oriented dialogue dataset with 11K dialogues and 83K emotion annotations of user utterances.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations