Evaluating Coherence in Dialogue Systems using Entailment (N19-1)

Copied to clipboard

Challenge: Evaluating open-domain dialogue systems is difficult due to the diversity of possible correct answers.
Approach: They propose a set of metrics for evaluating topic coherence using distributed sentence representations and calculable approximations of human judgment using conversational coherency.
Outcome: The proposed metrics can be used as a surrogate for human judgment based on conversational coherence on large-scale datasets and provide an unbiased estimate for the quality of the responses.

Similar Papers

GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems (2020.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics only consider surface features or utterance-level semantics, without explicitly considering the fine-grained topic transition dynamics of dialogue flows.
Approach: They propose a graph-enhanced evaluation metric GRADE to evaluate dialogue coherence . GRADE incorporates utterance-level contextualized representations and fine-grained topic-level graph representations to improve communication logic.
Outcome: The proposed evaluation metric outperforms state-of-the-art metrics on measuring diverse dialogue models in terms of Pearson and Spearman correlations with human judgments.
Towards Holistic and Automatic Evaluation of Open-Domain Dialogue Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods of open-domain dialogue evaluation are labor-intensive and inefficient.
Approach: They propose to use open-domain dialogues to evaluate different aspects of dialogues using holistic evaluation metrics.
Outcome: The proposed metrics show strong correlations with human judgments.
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems (2022.acl-long)

Copied to clipboard

Challenge: Evaluation of open-domain dialogue systems is challenging and unreliable . human evaluation of live conversations is highly reliable, but reliability cannot be assumed .
Approach: They propose a method of open-domain dialogue evaluation that is highly reliable . they compare live conversations with models that avoid pre-created reference dialogues .
Outcome: The proposed method is highly reliable while remaining feasible and low cost.
Dialogue Coherence Assessment Without Explicit Dialogue Act Labels (2020.acl-main)

Copied to clipboard

Challenge: Recent dialogue coherence models use coherency features designed for monologue texts to represent utterances and then explicitly augment them with dialogue-relevant features, e.g., dialogue act labels.
Approach: They propose a multi-task learning approach that uses dialogue act prediction to obtain informative utterance representations for coherence assessment.
Outcome: The proposed model outperforms its strong competitors on the DailyDialogue corpus and performs on par with them on the SwitchBoard corpus for ranking dialogues concerning their coherence.
Automating Human Evaluation of Dialogue Systems (2022.naacl-srw)

Copied to clipboard

Challenge: a recent study shows that human evaluations of dialogue systems weakly reflect human judgments.
Approach: They propose a BERT-based model that fine-tunes a model with three prediction heads to predict whether the system-generated output is natural, fluent, and informative.
Outcome: The proposed model achieves an average accuracy of 77% over the 3 labels . it also uses three different models to compute the labels compared to three separate models .
Evaluating Dialogue Generation Systems via Response Selection (2020.acl-main)

Copied to clipboard

Challenge: Existing automatic evaluation metrics for open-domain dialogue systems correlate poorly with human evaluation.
Approach: They propose to construct response selection test sets with well-chosen false candidates to evaluate response generation systems via response selection.
Outcome: The proposed method correlates with human evaluation better than widely used metrics such as BLEU.
Large-Scale Correlation Analysis of Automated Metrics for Topic Models (2023.acl-long)

Copied to clipboard

Challenge: Existing studies on topic models lack a correlation between automated coherence metrics and human judgement.
Approach: They propose a sampling approach to mine topics for metric evaluation and extend the analysis to measure topical differences between corpora.
Outcome: The proposed method extends to measure topical differences between corpora and human judgement by using extensive user study.
uBLEU: Uncertainty-Aware Automatic Evaluation Method for Open-Domain Dialogue Systems (2020.acl-srw)

Copied to clipboard

Challenge: Existing evaluation metrics for text generation tasks do not consider uncertain responses without writing additional reference responses by hand.
Approach: They propose a human-aided, uncertainty-aware evaluation method for open-domain dialogue systems, BLEU.
Outcome: The proposed method is comparable to existing methods on Twitter and improves state-of-the-art evaluation method RUBER.
Learning an Unreferenced Metric for Online Dialogue Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Existing tools for dialogue evaluation do not generalize to unseen datasets and/or need a human-generated reference response during inference.
Approach: They propose an unreferenced automated dialogue evaluation metric that uses large pre-trained language models to extract latent representations of utterances and leverages the temporal transitions that exist between them.
Outcome: The proposed model achieves higher correlation with human annotations in an online setting, while not requiring true responses for comparison during inference.
DialSummEval: Revisiting Summarization Evaluation for Dialogues (2022.naacl-main)

Copied to clipboard

Challenge: Current models for dialogue summarization have flaws that may not be well exposed by frequently used metrics such as ROUGE.
Approach: They propose to re-evaluate 18 categories of metrics in terms of four dimensions: coherence, consistency, fluency and relevance, as well as a unified human evaluation of various models for the first time.
Outcome: The proposed dataset will be used to evaluate 18 categories of metrics in terms of coherence, consistency, fluency and relevance, and a unified human evaluation of various models for the first time.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations