Challenge: Using fine-grained evaluation techniques, translation outputs have become better and more fluent.
Approach: They propose a fine-grained test suite for the language pair German–English . they describe the creation and implementation of the test suite in detail .
Outcome: The proposed test suite is based on linguistically motivated categories and phenomena and semi-automatic evaluation is carried out with regular expressions.

Similar Papers

TQ-AutoTest – An Automated Test Suite for (Machine) Translation Quality (L18-1)

Copied to clipboard

Challenge: Especially the trend towards neural MT has renewed peoples' interest in better and more analytical diagnostic methods for MT quality.
Approach: They propose a framework that supports a linguistic evaluation of machine translations using test suites.
Outcome: The proposed framework supports linguistic evaluation of (machine) translations using test suites.
A French Version of the FraCaS Test Suite (2020.lrec-1)

Copied to clipboard

Challenge: a French version of the FraCaS test suite is presented in this paper . it contains problems illustrating semantic inference in natural language .
Approach: They propose to test the NLP system's semantic capacity against inferencing tasks by translating the FraCaS test suite into French and running an experiment to test both the translation and the logical semantics underlying the problems.
Outcome: The proposed tests were compared with similar tests conducted in other languages and show that the results are comparable to those of other tests.
Informative Manual Evaluation of Machine Translation Output (2020.coling-main)

Copied to clipboard

Challenge: a new method for manual evaluation of machine translation output is proposed . evaluators mark problematic parts of the translated text, not just overall scores .
Approach: They propose a method for manual evaluation of machine translation output based on marking actual issues in the translated text.
Outcome: The proposed method can be applied on any genre/domain and language pair . it can be guided by various types of quality criteria and can be used for other types of generated text.
MTNT: A Testbed for Machine Translation of Noisy Text (D18-1)

Copied to clipboard

Challenge: Noisy input text can cause disastrous mistranslations in most modern machine translation systems.
Approach: They propose a benchmark dataset for Machine Translation of Noisy Text (MTNT) they use reddit comments and professionally sourced translations to examine noise types.
Outcome: The proposed dataset can provide an attractive testbed for noise-robust machine translation systems.
Evaluating Pronominal Anaphora in Machine Translation: An Evaluation Measure and a Test Suite (D19-1)

Copied to clipboard

Challenge: Currently, machine translation is performed at the level of individual sentences, in isolation from the rest of the document.
Approach: They propose a dataset that can be used as a test suite for pronoun translation . they propose an evaluation measure to differentiate good and bad pronounce translations .
Outcome: The proposed dataset can be used as a test suite for pronoun translation in English . it covers multiple source languages and different pronouner errors drawn from real system translations .
DOLFIN - Document-Level Financial Test-Set for Machine Translation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing document-level machine translation test-sets cover general domain but fall short on specialised domains, such as legal and financial.
Approach: They propose to use a document-level machine translation test-set to replace perfectly aligned sentences by presenting data in units of sections rather than sentences.
Outcome: The proposed dataset is built from specialised financial documents and it shows that it can discriminate between context-sensitive and context-agnostic models and shows the weaknesses when models fail to accurately translate financial texts.
Tilde MT Platform for Developing Client Specific MT Solutions (L18-1)

Copied to clipboard

Challenge: a growing demand for translations and multilingual content is surpassing the supply of professional translation services.
Approach: They present a custom machine translation platform called Tilde MT that provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality.
Outcome: The proposed platform provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality, and wide integration capabilities.
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on English and a handful of high-resource languages, thereby overlooking the realistic performance of large language models in multilingual and lower-resourced scenarios.
Approach: They propose a unified and lightweight framework that integrates 27 benchmarks under a standard ISO 639-3 language identifier system to enable seamless incorporation of new benchmarks.
Outcome: The proposed framework integrates 27 benchmarks under a standard ISO 639-3 language identifier system, allowing for seamless incorporation of new benchmarks.
On Context Span Needed for Machine Translation Evaluation (2020.lrec-1)

Copied to clipboard

Challenge: a number of common patterns can be observed for context-aware MT evaluation, authors say . document-level evaluations have largely been performed at the sentence level . the definition of what constitutes a "document level" evaluation is still unclear .
Approach: They propose to use a series of surveys to identify the necessary context span . they find common patterns that can be used to draw general guidelines .
Outcome: The proposed evaluations of machine translation systems show that some issues and spans depend on domain and target language.
Evaluating Automatic Metrics with Incremental Machine Translation Systems (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have shown that neural metrics are more reliable than non-neural metrics.
Approach: They propose to use commercial machine translations to evaluate machine translation metrics based on their preference for more recent outputs.
Outcome: The proposed dataset confirms several previous findings, including the advantage of neural metrics over non-neural ones, and also explores the debated issue of how MT quality affects metric reliability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations