Train, Sort, Explain: Learning to Diagnose Translation Models (N19-4)

Copied to clipboard

Challenge: Evaluating translation models is a trade-off between effort and detail.
Approach: They propose to use a neural text classifier to automatically expose systematic differences between human and machine translations to human experts.
Outcome: The proposed method exposes systematic differences between human and machine translations to human experts.

Similar Papers

Human or Neural Translation? (2020.coling-main)

Copied to clipboard

Challenge: a recent study shows that deep neural models have improved machine translation . identifying machine translation is still feasible, but is not yet known.
Approach: They train and apply deep neural models to distinguish between human and machine translations . they use a monolingual and bilingual task to train and train 18 classifiers based on their results .
Outcome: The proposed model improves the ability to distinguish between human and machine translations at the sentence level.
Towards Modeling the Style of Translators in Neural Machine Translation (2021.naacl-main)

Copied to clipboard

Challenge: a key ingredient of neural machine translation is the use of large datasets with different but consistent translation styles . however, the models do not capture the variety of translators' styles from the data . a recent study shows that style-augmented models can capture the style variations of translator .
Approach: They propose to augment a neural machine translation model with translator information . they use TED talk datasets to model and control translator-related stylistic variations .
Outcome: The proposed models capture the style variations of translators and generate translations with different styles on new data.
Lost in Translation, and Found: Detecting and Interpreting Translation Effects (2026.acl-long)

Copied to clipboard

Challenge: Translationese refers to the statistical patterns that distinguish translated texts from original texts.
Approach: They analyze linguistic features which enable our model to achieve high accuracy by a collection of linguistic characteristics and pretrained neural models pick up these features without any fine-tuning.
Outcome: The proposed model achieves high accuracy with a set of linguistic features that correspond to translationese theories and pretrained neural models pick up these features without any fine-tuning.
Revisit Automatic Error Detection for Wrong and Missing Translation – A Supervised Approach (D19-1)

Copied to clipboard

Challenge: Current machine translation techniques are bottlenecked by adequacy issues . we propose automatic detection of missing and wrong translations .
Approach: They propose automatic detection of adequacy errors in MT hypothesis for MT model evaluation by annotating missing and wrong translations in 15000 Chinese-English translation pairs.
Outcome: The proposed model can detect missing and wrong translations in 15000 Chinese-English translation pairs.
Why Generate When You Can Discriminate? A Novel Technique for Text Classification using Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Existing methods for text classification using autoregressive language models are limited . authors propose a novel technique for text classification using autoreregressives .
Approach: They propose a two-step technique for text classification using autoregressive language models . they use a set of perplexity and log-likelihood based numeric features to elicit a text instance .
Outcome: The proposed technique eliminates parameter updates in LMs and does not limit training examples . it is evaluated across 5 datasets and compares with multiple competent baselines .
A Discriminative Neural Model for Cross-Lingual Word Alignment (D19-1)

Copied to clipboard

Challenge: a novel word alignment model for machine translation has been developed for a number of languages . explicit word-to-word alignments have largely been lost in neural MT systems .
Approach: They propose a discriminative word alignment model which integrates into a Transformer-based machine translation model.
Outcome: The proposed model performs better on Chinese and Arabic alignments than standard models.
Online Learning Meets Machine Translation Evaluation: Finding the Best Systems with the Least Human Effort (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to evaluate multiple systems are expensive and require human evaluators.
Approach: They propose a novel online learning approach that dynamically converges to the top-3 ranked systems for the language pairs considered by taking advantage of human feedback.
Outcome: The proposed approach converges to the top-3 ranked systems for the language pairs considered despite the lack of human feedback for many translations.
Controlling Text Complexity in Neural Machine Translation (D19-1)

Copied to clipboard

Challenge: Prior work on text complexity has focused on simplifying input text in one language, primarily English.
Approach: They propose a method to align news articles written for different levels of target language proficiency.
Outcome: The proposed model outperforms pipeline approaches that translate and simplify text independently.
Discriminative Reranking for Neural Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: reranking models allow the integration of rich features to select a better output hypothesis within an n-best list or lattice.
Approach: They use discriminative reranking to train a large transformer architecture to train an ranked list of hypotheses.
Outcome: Experiments on four WMT directions show that discriminative reranking improves translation quality.
Comparing Feature-Engineering and Feature-Learning Approaches for Multilingual Translationese Classification (2021.emnlp-main)

Copied to clipboard

Challenge: Traditional hand-crafted features have been used for distinguishing between translated and original non-translated texts.
Approach: They compare a feature-engineering-based approach to a features-learning-based one and use pre-trained neural word embeddings to train neural architectures.
Outcome: The proposed approach outperforms other approaches by more than 20 accuracy points and the BERT-based model performs the best in both monolingual and multilingual settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations