Challenge: Sentence-level Quality estimation (QE) is traditionally a regression task . but large multilingual contextualized language models are expensive and infeasible for real-world applications.
Approach: They evaluate several model compression techniques for QE and find they are inefficient . they argue that a full model parameterization is required to achieve SoTA results .
Outcome: The proposed models are poorly expressive in a regression task, the authors argue . they show that reframing QE as a classification problem and evaluating models would improve their performance in real-world applications.

Similar Papers

Are we Estimating or Guesstimating Translation Quality? (2020.acl-main)

Copied to clipboard

Challenge: A carefully engineered ensemble of pre-trained multilingual language models won the QE shared task at WMT19.
Approach: They propose to use pre-trained multilingual language models to train quality estimation for machine translation.
Outcome: A carefully engineered ensemble of pre-trained language models wins the QE shared task at WMT19.
Assessing Quality Estimation Models for Sentence-Level Prediction (C18-1)

Copied to clipboard

Challenge: Using a relevant QE model is also very important in QE.
Approach: They evaluate a wide range of advanced sentence-level Quality Estimation models including Support Vector Regression, Ride Regression and Bayesian Neural Networks.
Outcome: The proposed models behave differently in evaluation settings depending on whether test data come from the same domain as the training data or not.
An Exploratory Analysis of Multilingual Word-Level Quality Estimation with Cross-Lingual Transformers (2021.acl-short)

Copied to clipboard

Challenge: Existing word-level quality estimation models require labelled data for each language pair and expensive maintenance.
Approach: They propose to use multilingual QE models to generalise across languages . they propose to train models on other language pairs to predict word-level quality .
Outcome: The proposed models generalise well across languages, making them more useful in real-world scenarios.
Self-Supervised Quality Estimation for Machine Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Training QE models require massive parallel data with hand-crafted quality annotations, which are time-consuming and labor-intensive to obtain.
Approach: They propose a self-supervised method to evaluate machine-translated sentences without references by recovering masked target words.
Outcome: The proposed method outperforms previous unsupervised methods on several QE tasks in different language pairs and domains.
Knowledge Distillation for Quality Estimation (2021.findings-acl)

Copied to clipboard

Challenge: Recent success in Quality Estimation stems from the use of multilingual pre-trained models, where large models lead to impressive results.
Approach: They propose to transfer knowledge from a strong QE teacher model to a much smaller model with a different, shallower architecture.
Outcome: The proposed model performs better than distilled models with 8x fewer parameters.
Translation Error Detection as Rationale Extraction (2022.findings-acl)

Copied to clipboard

Challenge: Recent Quality Estimation models rely on translation errors to predict overall sentence quality, but detecting specific errors is a more challenging task.
Approach: They propose to use a semi-supervised method to detect translation errors by attribution of relevance scores to inputs to explain model predictions.
Outcome: The proposed method can detect translation errors and is compared with human models using a set of feature attribution methods.
Rethinking the Word-level Quality Estimation for Machine Translation from Human Judgement (2023.findings-acl)

Copied to clipboard

Challenge: Word-level Quality Estimation (QE) of Machine Translation aims to detect potential translation errors in the translated sentence without reference.
Approach: They propose to use a human-generated translation judgment to generate a word-level quality estimate (QE) using a translation error rate toolkit to detect translation errors without reference.
Outcome: The proposed dataset is more consistent with human judgment and confirms the effectiveness of the proposed tag-correcting strategies.
A Multi-task Learning Framework for Quality Estimation (2023.findings-acl)

Copied to clipboard

Challenge: Conventional approaches to QE involve training separate models at different levels of granularity viz., word-level, sentence-level and document-level .
Approach: They propose to train a single model for sentence-level and word-level QE tasks in a multi-task learning framework and compare them to baseline models.
Outcome: The proposed model improves on the single-pair, multi-patch, and zero-shot settings.
Quality Estimation without Human-labeled Data (2021.eacl-main)

Copied to clipboard

Challenge: Quality estimation aims to measure the quality of translated content without access to a reference translation.
Approach: They propose a method that uses synthetic training data to train supervised quality estimation models.
Outcome: The proposed model outperforms models trained on human-annotated data for sentence and word-level prediction.
Unsupervised Quality Estimation for Neural Machine Translation (2020.tacl-1)

Copied to clipboard

Challenge: Existing approaches require large amounts of expert annotated data, computation, and time for training.
Approach: They propose an unsupervised approach to QE where no training is required . they use a dataset that enables work on both black-box and glass-box approaches .
Outcome: The proposed approach rivals state-of-the-art supervised QE models in terms of correlation with human judgments of quality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations