| Challenge: | Evaluating open-domain dialogue systems is difficult due to the diversity of possible correct answers. |
| Approach: | They propose a set of metrics for evaluating topic coherence using distributed sentence representations and calculable approximations of human judgment using conversational coherency. |
| Outcome: | The proposed metrics can be used as a surrogate for human judgment based on conversational coherence on large-scale datasets and provide an unbiased estimate for the quality of the responses. |
Similar Papers
GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics only consider surface features or utterance-level semantics, without explicitly considering the fine-grained topic transition dynamics of dialogue flows. |
| Approach: | They propose a graph-enhanced evaluation metric GRADE to evaluate dialogue coherence . GRADE incorporates utterance-level contextualized representations and fine-grained topic-level graph representations to improve communication logic. |
| Outcome: | The proposed evaluation metric outperforms state-of-the-art metrics on measuring diverse dialogue models in terms of Pearson and Spearman correlations with human judgments. |
Towards Holistic and Automatic Evaluation of Open-Domain Dialogue Generation (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods of open-domain dialogue evaluation are labor-intensive and inefficient. |
| Approach: | They propose to use open-domain dialogues to evaluate different aspects of dialogues using holistic evaluation metrics. |
| Outcome: | The proposed metrics show strong correlations with human judgments. |
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems (2022.acl-long)
Copied to clipboard
| Challenge: | Evaluation of open-domain dialogue systems is challenging and unreliable . human evaluation of live conversations is highly reliable, but reliability cannot be assumed . |
| Approach: | They propose a method of open-domain dialogue evaluation that is highly reliable . they compare live conversations with models that avoid pre-created reference dialogues . |
| Outcome: | The proposed method is highly reliable while remaining feasible and low cost. |
Dialogue Coherence Assessment Without Explicit Dialogue Act Labels (2020.acl-main)
Copied to clipboard
| Challenge: | Recent dialogue coherence models use coherency features designed for monologue texts to represent utterances and then explicitly augment them with dialogue-relevant features, e.g., dialogue act labels. |
| Approach: | They propose a multi-task learning approach that uses dialogue act prediction to obtain informative utterance representations for coherence assessment. |
| Outcome: | The proposed model outperforms its strong competitors on the DailyDialogue corpus and performs on par with them on the SwitchBoard corpus for ranking dialogues concerning their coherence. |
Automating Human Evaluation of Dialogue Systems (2022.naacl-srw)
Copied to clipboard
| Challenge: | a recent study shows that human evaluations of dialogue systems weakly reflect human judgments. |
| Approach: | They propose a BERT-based model that fine-tunes a model with three prediction heads to predict whether the system-generated output is natural, fluent, and informative. |
| Outcome: | The proposed model achieves an average accuracy of 77% over the 3 labels . it also uses three different models to compute the labels compared to three separate models . |
Evaluating Dialogue Generation Systems via Response Selection (2020.acl-main)
Copied to clipboard
| Challenge: | Existing automatic evaluation metrics for open-domain dialogue systems correlate poorly with human evaluation. |
| Approach: | They propose to construct response selection test sets with well-chosen false candidates to evaluate response generation systems via response selection. |
| Outcome: | The proposed method correlates with human evaluation better than widely used metrics such as BLEU. |
Large-Scale Correlation Analysis of Automated Metrics for Topic Models (2023.acl-long)
Copied to clipboard
| Challenge: | Existing studies on topic models lack a correlation between automated coherence metrics and human judgement. |
| Approach: | They propose a sampling approach to mine topics for metric evaluation and extend the analysis to measure topical differences between corpora. |
| Outcome: | The proposed method extends to measure topical differences between corpora and human judgement by using extensive user study. |
uBLEU: Uncertainty-Aware Automatic Evaluation Method for Open-Domain Dialogue Systems (2020.acl-srw)
Copied to clipboard
| Challenge: | Existing evaluation metrics for text generation tasks do not consider uncertain responses without writing additional reference responses by hand. |
| Approach: | They propose a human-aided, uncertainty-aware evaluation method for open-domain dialogue systems, BLEU. |
| Outcome: | The proposed method is comparable to existing methods on Twitter and improves state-of-the-art evaluation method RUBER. |
Learning an Unreferenced Metric for Online Dialogue Evaluation (2020.acl-main)
Copied to clipboard
| Challenge: | Existing tools for dialogue evaluation do not generalize to unseen datasets and/or need a human-generated reference response during inference. |
| Approach: | They propose an unreferenced automated dialogue evaluation metric that uses large pre-trained language models to extract latent representations of utterances and leverages the temporal transitions that exist between them. |
| Outcome: | The proposed model achieves higher correlation with human annotations in an online setting, while not requiring true responses for comparison during inference. |
DialSummEval: Revisiting Summarization Evaluation for Dialogues (2022.naacl-main)
Copied to clipboard
| Challenge: | Current models for dialogue summarization have flaws that may not be well exposed by frequently used metrics such as ROUGE. |
| Approach: | They propose to re-evaluate 18 categories of metrics in terms of four dimensions: coherence, consistency, fluency and relevance, as well as a unified human evaluation of various models for the first time. |
| Outcome: | The proposed dataset will be used to evaluate 18 categories of metrics in terms of coherence, consistency, fluency and relevance, and a unified human evaluation of various models for the first time. |