Challenge: Existing methods to measure diversity of chitchat model responses have been proposed to measure iteratively.
Approach: They propose a metric which uses Natural Language Inference to measure the semantic diversity of a set of model responses for a conversation.
Outcome: The proposed metric improves the diversity of a sampled set of responses using a new generation procedure called Diversity Threshold Generation.

Similar Papers

Measuring and Improving Semantic Diversity of Dialogue Generation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation metrics for response diversity do not capture the semantic diversity of generated responses.
Approach: They propose an automatic evaluation metric to measure the semantic diversity of generated responses . they show that it captures human judgments better than existing diversity metrics .
Outcome: The proposed metric captures human judgments on response diversity better than existing lexical diversity metrics.
Dialogue Natural Language Inference (P19-1)

Copied to clipboard

Challenge: Consistency is a long standing issue faced by dialogue models.
Approach: They propose to frame the consistency of dialogue agents as natural language inference and create a new natural language dataset called Dialogue NLI.
Outcome: The proposed model can improve the consistency of a dialogue model with human evaluation and automatic metrics on a suite of evaluation sets designed to measure the model’s consistency.
More Diverse Dialogue Datasets via Diversity-Informed Data Collection (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to generate conversational dialogue produce uninteresting, predictable responses.
Approach: They propose a method to collect and determine more diverse data from conversational participants . they use dynamically computed corpus-level statistics to determine which conversational participant to collect data from .
Outcome: The proposed method produces significantly more diverse data than baseline methods and better results on emotion classification and dialogue generation tasks.
Evaluating the Evaluation of Diversity in Natural Language Generation (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for controlling diversity by tuning a “decoding parameter” affect form but not meaning.
Approach: They propose a framework that measures correlation between a diversity metric and a parameter that controls some aspect of diversity in generated text.
Outcome: The proposed framework outperforms existing methods in estimating diversity . it shows that humans outperformed existing methods but affect form but not meaning .
Diversifying Dialogue Generation with Non-Conversational Text (2020.acl-main)

Copied to clipboard

Challenge: Neural network-based sequence-to-sequence models suffer from low diversity in open-domain dialogue generation.
Approach: They propose a way to diversify dialogue generation by leveraging non-conversational text . they collect large-scale corpus from forum comments, idioms and book snippets .
Outcome: The proposed model produces significantly more diverse responses without sacrificing relevance with context.
Semantic Diversity for Natural Language Understanding Evaluation in Dialog Systems (2020.coling-industry)

Copied to clipboard

Challenge: a dialog system is used to evaluate NLU models using aggregated metrics on a large number of utterances.
Approach: They propose a method to generate a test set with high semantic diversity for NLU evaluation in dialog systems.
Outcome: The proposed test sets are based on high diversity of utterances from different regions of the utteration embedding space.
Towards Holistic and Automatic Evaluation of Open-Domain Dialogue Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods of open-domain dialogue evaluation are labor-intensive and inefficient.
Approach: They propose to use open-domain dialogues to evaluate different aspects of dialogues using holistic evaluation metrics.
Outcome: The proposed metrics show strong correlations with human judgments.
Better Conversations by Modeling, Filtering, and Optimizing for Coherence and Diversity (D18-1)

Copied to clipboard

Challenge: Existing encoder-decoder models for open domain dialogue generate generic, uninformative, and non-coherent responses.
Approach: They propose to introduce a measure of coherence as the GloVe embedding similarity between dialogue context and generated response to improve output diversity.
Outcome: The proposed model improves on the OpenSubtitles corpus in terms of BLEU score and diversity metrics.
Informed Sampling for Diversity in Concept-to-Text NLG (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to encourage lexical diversity for language generation tasks produce repetitive outputs, but this often comes at a cost to the perceived fluency and adequacy of the output.
Approach: They propose to augment the decoding process with a meta-classifier trained to distinguish which words at any given timestep will lead to high-quality output.
Outcome: The proposed method achieves a high level of diversity with minimal effect on the output’s fluency and adequacy.
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code (2025.findings-emnlp)

Copied to clipboard

Challenge: Language models (LMs) have exhibited impressive abilities in generating code from natural language requirements.
Approach: They propose to introduce various metrics with inter-code similarity to evaluate the diversity of generated code by comparing model-generated solutions with human-written ones.
Outcome: The proposed method leverages LMs’ capabilities in code understanding and reasoning, resulting in a set of metrics that represent the number of algorithms in model-generated solutions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations