Challenge: Existing evaluation metrics focus on turnlevel quality, which is not well suited for open-end dialogue tasks.
Approach: They propose to measure the performance of a dialogue system by computing the distributionwise distance between its generated conversations and real-world conversations.
Outcome: The proposed metrics correlate better with human judgments than existing metrics on dialogue systems.

Similar Papers

Don’t Forget Your ABC’s: Evaluating the State-of-the-Art in Chat-Oriented Dialogue Systems (2023.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods are biased because of their subjectivity and inconsistent evaluation can misinform the performance of a chat-oriented open-domain dialogue system.
Approach: They propose to use a human evaluation method to estimate the rates of manypasted macro ‘LN’ dialogue system behaviors to compare them with existing evaluation methods.
Outcome: The proposed method is more suitable than alternative Likert-style or comparative approaches for dimensional evaluation of open-domain dialogue systems.
Towards Holistic and Automatic Evaluation of Open-Domain Dialogue Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods of open-domain dialogue evaluation are labor-intensive and inefficient.
Approach: They propose to use open-domain dialogues to evaluate different aspects of dialogues using holistic evaluation metrics.
Outcome: The proposed metrics show strong correlations with human judgments.
Evaluating Dialogue Generation Systems via Response Selection (2020.acl-main)

Copied to clipboard

Challenge: Existing automatic evaluation metrics for open-domain dialogue systems correlate poorly with human evaluation.
Approach: They propose to construct response selection test sets with well-chosen false candidates to evaluate response generation systems via response selection.
Outcome: The proposed method correlates with human evaluation better than widely used metrics such as BLEU.
DiQAD: A Benchmark Dataset for Open-domain Dialogue Quality Assessment (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on dialogue quality assessment are uncapable of providing an end-to-end and human-epistemic assessment dataset . open-domain dialogue assessment is complicated and costly, but it can be done by recruiting human evaluators.
Approach: They propose a large-scale dialogue quality assessment dataset for automatically assessing open-domain dialogue quality.
Outcome: The proposed dataset is openly accessible at https://github.com/yukunZhao/Dialogue_quality_evaluation.
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems (2022.acl-long)

Copied to clipboard

Challenge: Evaluation of open-domain dialogue systems is challenging and unreliable . human evaluation of live conversations is highly reliable, but reliability cannot be assumed .
Approach: They propose a method of open-domain dialogue evaluation that is highly reliable . they compare live conversations with models that avoid pre-created reference dialogues .
Outcome: The proposed method is highly reliable while remaining feasible and low cost.
Learning an Unreferenced Metric for Online Dialogue Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Existing tools for dialogue evaluation do not generalize to unseen datasets and/or need a human-generated reference response during inference.
Approach: They propose an unreferenced automated dialogue evaluation metric that uses large pre-trained language models to extract latent representations of utterances and leverages the temporal transitions that exist between them.
Outcome: The proposed model achieves higher correlation with human annotations in an online setting, while not requiring true responses for comparison during inference.
Evaluating Coherence in Dialogue Systems using Entailment (N19-1)

Copied to clipboard

Challenge: Evaluating open-domain dialogue systems is difficult due to the diversity of possible correct answers.
Approach: They propose a set of metrics for evaluating topic coherence using distributed sentence representations and calculable approximations of human judgment using conversational coherency.
Outcome: The proposed metrics can be used as a surrogate for human judgment based on conversational coherence on large-scale datasets and provide an unbiased estimate for the quality of the responses.
FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: Recent model-based reference-free metrics for open-domain dialogue evaluation lack correlations with human judgment and poor interpretability.
Approach: They propose a multi-dimensional dialogue-level metric with three sub-metrics targeting a specific dimension.
Outcome: The proposed metric outperforms existing models and sub-metrics in three high-quality dialogue evaluation benchmarks.
CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems (2025.coling-main)

Copied to clipboard

Challenge: Existing metrics for dialogue quality evaluation show low correlation with human judgements . current metrics do not accurately evaluate dialogue responses based on dialogue history .
Approach: They propose a new metric measuring causal strength between dialogue histories and responses . they collect a dialogue dataset with human-annotated causal relations and pairwise human judgements .
Outcome: The proposed metric outperforms existing state-of-the-art metrics in human judgements . it is based on a dialogue dataset with human-annotated causal relations and human judgement sets .
USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation (2020.acl-main)

Copied to clipboard

Challenge: Standard language generation metrics have been shown to be ineffective for dialog evaluation.
Approach: They propose an unsupervised evaluation metric for dialog that trains unsupervised models to measure several desirable qualities of dialog.
Outcome: The proposed evaluation metric strongly correlates with human judgment on Topical-Chat and PersonaChat.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations