Challenge: Existing approaches to training or evaluating non-English dialogue datasets often introduce artifacts that reduce their naturalness and cultural appropriateness.
Approach: They propose a structured framework for encoding, localizing, and generating multilingual dialogues from abstract intent representations.
Outcome: The proposed framework outperforms translation models in Italian, German, and Chinese on cultural relevance, coherence, and situational appropriateness.

Similar Papers

Cross-Lingual Dialogue Dataset Creation via Outline-Based Generation (2023.tacl-1)

Copied to clipboard

Challenge: Multilingual task-oriented dialogue (ToD) datasets suffer from severe limitations, such as being small in scale and lacking naturalness and cultural specificity in the target language.
Approach: They propose a novel outline-based annotation process where domain-specific abstract schemata of dialogue are mapped into natural language outlines.
Outcome: The proposed approach improves understanding, dialogue state tracking, and end-to-end dialogue evaluation in Arabic, Indonesian, Russian, and Kiswahili.
Natural Language Processing for Multilingual Task-Oriented Dialogue (2022.acl-tutorials)

Copied to clipboard

Challenge: a tutorial will examine the challenges and gaps in multilingual ToD research . multilingual systems are difficult to build, and are limited to English and other languages .
Approach: This tutorial will discuss the importance of multilingual task-oriented dialogue systems . it will provide an overview of current research gaps, challenges and initiatives related to multilingual ToD systems - with a particular focus on their connections to current research and challenges in multilingual and low-resource NLP.
Outcome: This tutorial will provide an overview of current research gaps, challenges and initiatives related to multilingual ToD systems.
Multi-Domain Dialogue Acts and Response Co-Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing pipeline approaches for task-oriented dialogue systems tend to predict multiple dialogue acts first and use them to assist response generation.
Approach: They propose a neural co-generation model that generates dialogue acts and responses concurrently and preserves semantic structures of multi-domain dialogue acts.
Outcome: The proposed model improves over state-of-the-art models in automatic and human evaluations on a large-scale dataset.
Multi-task Learning for Natural Language Generation in Task-Oriented Dialogue (D19-1)

Copied to clipboard

Challenge: Existing methods to generate natural language for task-oriented dialogues lack naturalness and variation in language.
Approach: They propose a multi-task learning framework for natural language generation that explicitly targets for naturalness in generated responses via an unconditioned language model.
Outcome: The proposed framework outperforms existing models across multiple datasets in the study of natural language generation.
Generalizable and Explainable Dialogue Generation via Explicit Action Learning (2020.findings-emnlp)

Copied to clipboard

Challenge: Conditioned response generation for task-oriented dialogues implicitly optimizes task completion and language quality.
Approach: They propose to learn natural language actions that represent utterances as a span of words.
Outcome: The proposed approach outperforms latent action baselines on a multi-domain dataset.
In Search of the Lost Arch in Dialogue: A Dependency Dialogue Acts Corpus for Multi-Party Dialogues (2025.findings-acl)

Copied to clipboard

Challenge: Understanding speaker intentions remains a challenge in NLP . a number of corpora annotated using theoretical frameworks of dialogue focus on utterance-level labeling of speaker intent, missing wider context, or the rhetorical structure of a dialogue.
Approach: They propose to annotate a corpus of 33 dialogues and over 9,000 utterance units using the Dependency Dialogue Acts framework.
Outcome: The proposed corpus spans four genres of multi-party conversations from different modalities.
Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations (2024.lrec-main)

Copied to clipboard

Challenge: Personalization is a multifaceted process that requires multiple definitions and varies between individuals.
Approach: They propose to systemically survey the recent landscape of personalized dialogue generation including the datasets employed, methodologies developed, and evaluation metrics applied.
Outcome: The proposed model can generate fluent and coherent responses to human queries in a language-based conversational agent.
Towards Continuous Dialogue Corpus Creation: writing to corpus and generating from it (L18-1)

Copied to clipboard

Challenge: Existing methods to create dialogue corpora annotated with interoperable semantic information are based on ISO standard data models and tools.
Approach: They propose to use a corpus as a shared repository for analysis and modelling of interactive dialogue behaviour and for implementation, integration and evaluation of dialogue system components.
Outcome: The proposed method is applied to the design of two multimodal interactive applications - the Virtual Negotiation Coach and the Virtual Debate Coach.
MulZDG: Multilingual Code-Switching Framework for Zero-shot Dialogue Generation (2022.coling-1)

Copied to clipboard

Challenge: Existing zero-shot dialogue generation systems rely on large-scale pre-trained language models.
Approach: They propose a multilingual learning framework for zero-shot dialogue generation that can transfer knowledge from an English corpus to a non-English corpus with zero samples.
Outcome: The proposed framework can transfer knowledge from an English corpus to a non-English corpus with zero samples.
Speaker Turn Modeling for Dialogue Act Classification (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to DA classification model utterances without incorporating the turn changes among speakers throughout the dialogue, thus treating it no different than non-interactive written text.
Approach: They propose to integrate the turn changes in conversations among speakers when modeling DAs by learning conversation-invariant speaker turn embeddings to represent speaker turns in a conversation.
Outcome: The proposed model captures semantics from the dialogue content while accounting for different speaker turns in a conversation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations