Challenge: Existing methods do not correlate strongly with human annotations.
Approach: They propose a method that measures the probability that a language model will continue the conversation with a fixed set of follow-ups.
Outcome: The proposed method achieves the highest correlation with human evaluations when compared against twelve existing methods.

Similar Papers

What is wrong with you?: Leveraging User Sentiment for Automatic Dialog Evaluation (2022.findings-acl)

Copied to clipboard

Challenge: Existing metrics for dialog evaluation are trained on human annotations, which is cumbersome to collect.
Approach: They propose to use user sentiment and other information as proxy to measure the quality of previous dialogs.
Outcome: The proposed model is comparable to models trained on human annotated data.
Designing Precise and Robust Dialogue Response Evaluators (2020.acl-main)

Copied to clipboard

Challenge: Existing automated dialogue response evaluators have only moderate correlation with human judgement and are not robust.
Approach: They propose to build a reference-free dialogue response evaluator that exploits the power of semi-supervised training and pretrained (masked) language models.
Outcome: The proposed model achieves strong correlation with human judgement and generalizes robustly to diverse responses and corpora.
Learning an Unreferenced Metric for Online Dialogue Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Existing tools for dialogue evaluation do not generalize to unseen datasets and/or need a human-generated reference response during inference.
Approach: They propose an unreferenced automated dialogue evaluation metric that uses large pre-trained language models to extract latent representations of utterances and leverages the temporal transitions that exist between them.
Outcome: The proposed model achieves higher correlation with human annotations in an online setting, while not requiring true responses for comparison during inference.
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems (2022.acl-long)

Copied to clipboard

Challenge: Evaluation of open-domain dialogue systems is challenging and unreliable . human evaluation of live conversations is highly reliable, but reliability cannot be assumed .
Approach: They propose a method of open-domain dialogue evaluation that is highly reliable . they compare live conversations with models that avoid pre-created reference dialogues .
Outcome: The proposed method is highly reliable while remaining feasible and low cost.
Evaluating Open-Domain Dialogues in Latent Space with Next Sentence Prediction and Mutual Information (2023.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods for open-domain dialogues are difficult due to the one-to-many issue of the open- domain dialogues.
Approach: They propose a learning-based automatic evaluation metric which can robustly evaluate open-domain dialogues by augmenting CVAEs with a Next Sentence Prediction objective and employing Mutual Information to model the semantic similarity of text in the latent space.
Outcome: The proposed method can evaluate open-domain dialogues on two open- domain dialogue datasets.
Improving Automated Evaluation of Open Domain Dialog via Diverse Reference Augmentation (2021.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that having multiple valid references is important for automated evaluations.
Approach: They propose a technique for automatically expanding a human generated reference to a set of candidate references.
Outcome: The proposed method improves correlations between human-generated metrics and human ratings of system outputs.
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Current evaluation practices of open domain dialogue systems are still highly dependent on human evaluation.
Approach: They propose to use an annotated dataset to evaluate chatbots using large language models.
Outcome: The proposed model improves over few-shot inferences on a GPT-3.5 generated dialogue dataset.
REAM♯: An Enhancement Approach to Reference-based Evaluation Metrics for Open-domain Dialog Generation (2021.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for open-domain dialogue systems are limited by the diversity of possible outcomings.
Approach: They propose a method to augment a reference set to improve reliability . they propose BLEU to measure similarity between a predicted response and a small set of references .
Outcome: The proposed model improves the reliability of reference-based metrics with augmented reference sets.
RADE: Reference-Assisted Dialogue Evaluation for Open-Domain Dialogue (2023.acl-long)

Copied to clipboard

Challenge: Evaluating open-domain dialogue systems is challenging because of the one-to-many problem.
Approach: They propose a reference-based dialogue evaluation approach that leverages the pre-created utterance as reference other than the gold response to relieve the one-to-many problem.
Outcome: The proposed method outperforms state-of-the-art evaluation methods on three datasets and two existing benchmarks.
xDial-Eval: A Multilingual Open-Domain Dialogue Evaluation Benchmark (2023.findings-emnlp)

Copied to clipboard

Challenge: Currently, human evaluation is the most reliable way to holistically judge the quality of the dialogue.
Approach: They propose to use English dialogue evaluation metrics to generalize them to other languages.
Outcome: The proposed metrics outperform OpenAI’s ChatGPT in terms of average Pearson correlations over all datasets and languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations