Challenge: Existing methods to create dialogue corpora annotated with interoperable semantic information are based on ISO standard data models and tools.
Approach: They propose to use a corpus as a shared repository for analysis and modelling of interactive dialogue behaviour and for implementation, integration and evaluation of dialogue system components.
Outcome: The proposed method is applied to the design of two multimodal interactive applications - the Virtual Negotiation Coach and the Virtual Debate Coach.

Similar Papers

A Brief Survey of Textual Dialogue Corpora (2022.lrec-1)

Copied to clipboard

Challenge: Several dialogue corpora are available for research purposes, but they do not cover all the necessities of real-world applications.
Approach: They analyze available dialogue corpora and propose possible approaches to create new ones.
Outcome: The proposed corpus of human-human dialogues is based on a list of available dialogue corpora . it covers speakers, size, languages, collection, annotations, and domains . some trends are identified and possible approaches are also discussed .
Multilingual Dialogue Generation and Localization with Dialogue Act Scripting (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to training or evaluating non-English dialogue datasets often introduce artifacts that reduce their naturalness and cultural appropriateness.
Approach: They propose a structured framework for encoding, localizing, and generating multilingual dialogues from abstract intent representations.
Outcome: The proposed framework outperforms translation models in Italian, German, and Chinese on cultural relevance, coherence, and situational appropriateness.
Dialogue Act Annotation in a Multimodal Corpus of First Encounter Dialogues (2020.lrec-1)

Copied to clipboard

Challenge: a method used to annotate dialogue acts in a multimodal corpus is described . the annotations allow for analysis of how multimodal signals contribute to the structure and content of the dialogues.
Approach: They propose to annotate dialogue acts in a multimodal corpus of first encounter dialogues . they focus on which dialogue acts often follow each other across speakers and which overlap gestural behaviour .
Outcome: The method used to annotate dialogue acts in a multimodal corpus is described.
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information.
Approach: They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants.
Outcome: The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants.
A French Medical Conversations Corpus Annotated for a Virtual Patient Dialogue System (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for creating virtual patient dialogue systems require large data specific to the language, domain and clinical cases studied.
Approach: They propose to build an annotated corpus of medical dialogues in french using medical interviews and a data annotation scheme.
Outcome: The proposed corpus is made publicly available under a Free/Libre Open Source licence.
Alignment Annotation for Clinic Visit Dialogue to Clinical Note Sentence Language Generation (2020.lrec-1)

Copied to clipboard

Challenge: Despite advances in natural language processing, converting a clinic visit conversation into a clinical note is a largely unexplored area of research.
Approach: They propose an annotation methodology that is content- and technique- agnostic while associating note sentences to sets of dialogue sentences.
Outcome: The proposed method is content- and technique-agnostic while associating note sentences to sets of dialogue sentences.
InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model (2025.findings-emnlp)

Copied to clipboard

Challenge: Spoken Dialogue models face challenges in handling nuanced interactional phenomena, such as interruptions and backchannels.
Approach: They propose to use a 150-hour English speech interaction dialogue dataset to empower spoken dialogue models with nuanced real-time interaction capabilities.
Outcome: The proposed dataset trains and evaluates a speech understanding model that classifies key interactional events directly from audio.
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization (D19-54)

Copied to clipboard

Challenge: Existing work on abstractive dialogue summarizations has focused on news summarizing but there is no such comprehensive dataset.
Approach: They propose to use a chat-dialogues corpus with abstractive dialogue summaries to generate a short version of text that covers the main points succinctly.
Outcome: The proposed dataset achieves higher ROUGE scores than the model-generated summaries of news, compared with human evaluators' judgement.
Automatic Generation of Large-scale Multi-turn Dialogues from Reddit (2022.coling-1)

Copied to clipboard

Challenge: Using a set of algorithms, we can generate large dialogue corpus from Reddit.
Approach: They propose to automatically convert posts and their comments from discussion forums such as Reddit into multi-turn dialogues.
Outcome: The proposed methods improve on the baseline method by 36.3% . the best method shows an improvement of 36.6% over the previous one .
A Synthetic Data Generation Framework for Grounded Dialogues (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to train grounded dialogues require large amounts of data.
Approach: They propose a synthetic data generation framework for grounded dialogues that takes knowledge data and heuristics to determine a dialogue flow and incrementally turn it into a dialog.
Outcome: The proposed framework significantly boosts model performance in training data and low-resource scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations