Challenge: Existing work assumes the Gaussian priors of the latent variable, which are incapable of representing complex latent variables effectively.
Approach: They propose to use the Dirichlet distribution with flexible structures to characterize latent variables in place of the Gaussian priors.
Outcome: The proposed model outperforms existing models on the dialogue generation task.

Similar Papers

A Hierarchical Latent Structure for Variational Conversation Modeling (N18-1)

Copied to clipboard

Challenge: Variational autoencoders suffer from the notorious degeneration problem, according to a new study . utterance drop regularization is an important feature of the hierarchical RNNs .
Approach: They propose a variational hierarchical conversation RNN framework that exploits latent variables and an utterance drop regularization to exploit latent variable.
Outcome: The proposed model outperforms state-of-the-art models on Cornell Movie Dialog and Ubuntu Dialog Corpus.
DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation (2022.acl-long)

Copied to clipboard

Challenge: Existing pre-trained dialog models shed light on various downstream tasks in natural language processing (NLP).
Approach: They propose a dialog pre-training framework that introduces latent variables into the enhanced encoder-decoder pre-train framework to increase relevance and diversity of responses.
Outcome: The proposed model achieves state-of-the-art on personaChat, DailyDialog, and DSTC7-AVSD datasets.
Variational Autoregressive Decoder for Neural Response Generation (D18-1)

Copied to clipboard

Challenge: Existing variational Bayesian models generate responses from a single latent variable, which is not sufficient to model high variability in responses.
Approach: They propose a conditional variable auto-encoder that sequentially introduces latent variables to condition the generation of each word in the response sequence.
Outcome: Empirical results show that the proposed model improves on state-of-the-art models on Opensubtitle and Reddit datasets.
Hierarchy Response Learning for Neural Conversation Generation (D19-1)

Copied to clipboard

Challenge: Neural conversation generation models can't perceive and express the intention effectively, causing dull and generic responses.
Approach: They propose a hierarchical response generation model to capture conversation intention . they propose an expression reconstruction model and an expression attention model .
Outcome: The proposed model can generate the responses with more appropriate content and expression.
RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to personalized dialogue generation rely on dialogue data paired with user traits, profiles or persona description sentences.
Approach: They propose a hierarchical transformer retriever trained on dialogue domain data to perform personalized retrieval and a context-aware prefix encoder that fuses the retrieved information to the decoder more effectively.
Outcome: The proposed model generates more fluent and personalized responses under a suite of human and automatic metrics and is superior to state-of-the-art baselines on English Reddit conversations.
Speculative Sampling in Variational Autoencoders for Dialogue Response Generation (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have tried to improve variational models but they fail to learn proper mappings.
Approach: They propose to use a variable-based sampling technique to find the most probable one from redundantly sampled latent variables to tie up the variable with a given response.
Outcome: The proposed method is effective in response generation with massive dialogue data constructed from Twitter posts.
Better Exploiting Latent Variables in Text Modeling (P19-1)

Copied to clipboard

Challenge: Consistent gains in performance on two datasets, Penn Treebank and Yahoo, indicate the generalizability of our method.
Approach: They propose a method to exploit latent variables through hidden state averaging by sampling latent variable multiple times at a gradient step.
Outcome: The proposed method shows consistent gains on two datasets showing that it is generalizable.
HL-EncDec: A Hybrid-Level Encoder-Decoder for Neural Response Generation (C18-1)

Copied to clipboard

Challenge: Existing models for conversation systems operate sentences at word-level . word-based models suffer from Unknown Words Issue and Preference Issue .
Approach: They propose a hybrid-level Encoder-Decoder model which utilizes word-level features and character-level ones.
Outcome: The proposed model outperforms non-word-level models in automatic metrics and human annotations on a Chinese corpus.
Recurrence Boosts Diversity! Revisiting Recurrent Latent Variable in Transformer-Based Variational AutoEncoder for Diverse Text Generation (2022.findings-emnlp)

Copied to clipboard

Challenge: Variational Auto-Encoder (VAE) has been widely adopted in text generation due to its ability to learn flexible representations.
Approach: They propose a Transformer-based recurrent VAE structure that imposes recurrence on segment-wise latent variables with arbitrarily separated text segments and constructs the posterior distribution with residual parameterization.
Outcome: The proposed structure can deduce a non-zero lower bound of the KL term and enhance the entanglement of each segment and preceding latent variables, providing a theoretical guarantee of generation diversity.
Implicit Deep Latent Variable Models for Text Generation (D19-1)

Copied to clipboard

Challenge: Variational auto-encoders have been used for text generation but their representation power is limited due to two reasons.
Approach: They advocate sample-based representations of variational distributions for natural language . they further develop an LVM to directly match the aggregated posterior to the prior .
Outcome: The proposed model can be viewed as a natural extension of VAEs with a regularization of maximizing mutual information, mitigating the "posterior collapse" issue.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations