Challenge: Existing methods to evaluate machine translation output are based on comparing MT output to one or more reference translations.
Approach: They propose to use probabilities given by a large, multilingual model as a reference-free metric.
Outcome: The proposed model is robust and likely to offer reasonable performance across a broad spectrum of domains and different system qualities.

Similar Papers

A Holistic Approach to Reference-Free Evaluation of Machine Translation (2023.acl-short)

Copied to clipboard

Challenge: Traditional machine translation evaluation relies on reference written by humans . reference-free evaluation gets rid of labor-intensive annotations, which can pivot easily to new domains .
Approach: They propose a reference-free evaluation approach that characterizes evaluation as two aspects: fluency and faithfulness.
Outcome: The proposed approach outperforms SOTA reference-fee metrics on machine translation datasets.
On the Limitations of Reference-Free Evaluations of Generated Text (2022.emnlp-main)

Copied to clipboard

Challenge: a recent study has shown that evaluation metrics which accurately estimate the quality of generated text are limited in their ability to evaluate generated text.
Approach: They argue that reference-free metrics are limited in their ability to evaluate generated text . they recommend that they be used as diagnostic tools for analyzing and understanding model behavior .
Outcome: The proposed evaluation metrics are limited in their ability to evaluate generated text . they can be optimized at test time, can be biased against models with similar outputs .
On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation Evaluation (2020.acl-main)

Copied to clipboard

Challenge: a standard evaluation setup for supervised machine learning tasks does not hold for natural language generation tasks.
Approach: They propose to use reference-free machine translation evaluation to compare source texts to system translations to find key limitations.
Outcome: The proposed metrics perform poorly as semantic encoders for reference-free machine translation evaluation.
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for judging metrics are sensitive to the translations used for evaluation, leading to falsely confident conclusions about a metric’s efficacy.
Approach: They propose a method for thresholding performance improvement under an automatic metric against human judgements by using a pairwise system ranking method.
Outcome: The proposed method allows quantification of type I versus type II errors incurred, i.e., insignificant human differences in system quality that are accepted, and significant human differences that are rejected.
Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 Papers (2021.acl-long)

Copied to clipboard

Challenge: a meta-evaluation of machine translation (MT) has been conducted in 769 research papers . a recent study shows that evaluation practices have changed over the past decade .
Approach: They propose a meta-evaluation method for machine translation that uses BLEU scores to evaluate MT performance.
Outcome: The proposed meta-evaluation of machine translation shows that evaluation practices have changed over the past decade . the authors suggest that the evaluation process should be streamlined and standardized to ensure the validity of the evaluation method .
Online Learning Meets Machine Translation Evaluation: Finding the Best Systems with the Least Human Effort (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to evaluate multiple systems are expensive and require human evaluators.
Approach: They propose a novel online learning approach that dynamically converges to the top-3 ranked systems for the language pairs considered by taking advantage of human feedback.
Outcome: The proposed approach converges to the top-3 ranked systems for the language pairs considered despite the lack of human feedback for many translations.
KoBE: Knowledge-Based Machine Translation Evaluation (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for machine translation evaluation do not require reference translations.
Approach: They propose a method for machine translation evaluation which does not require reference translations.
Outcome: The proposed method achieves highest correlation with human judgements on 9 out of 18 language pairs from the WMT19 benchmark for evaluation without references.
Spurious Correlations in Reference-Free Evaluation of Text Generation (2022.acl-long)

Copied to clipboard

Challenge: Recent work suggests that reference-free evaluation metrics may rely on spurious correlations with human judgments.
Approach: They propose to use model-based, reference-free evaluation metrics to evaluate natural language generation systems.
Outcome: The proposed metrics achieve high correlations with human judgments, but they may not be robust enough to evaluate their efficacy and robustness.
BLEU might be Guilty but References are not Innocent (2020.emnlp-main)

Copied to clipboard

Challenge: Using a method to collect references and compare their value with human evaluations, we show that multi-reference BLEU does not improve the correlation for high quality output.
Approach: They propose a method to compare the quality of automated metrics by analyzing references and comparing them with human evaluations.
Outcome: The proposed method improves correlation with all modern evaluation metrics including embedding-based methods.
Is Reference Necessary in the Evaluation of NLG Systems? When and Where? (2024.naacl-long)

Copied to clipboard

Challenge: Despite recent advances in reference-free metrics, it has not been well understood when and where they can be used as an alternative to reference-based metrics.
Approach: They propose to use reference-free metrics to evaluate NLG systems . they find they have a higher correlation with human judgment and greater sensitivity to deficiencies in language quality .
Outcome: The proposed metrics exhibit higher correlation with human judgment and greater sensitivity to deficiencies in language quality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations