Papers by Christian Federmann
Multilingual Whispers: Generating Paraphrases with Translation (D19-55)
Copied to clipboard
| Challenge: | Humans naturally paraphrase, but they can generate approximately the same meaning with a different surface realization. |
| Approach: | They compare translation-based paraphrase gathering using human, automatic, or hybrid techniques to monolingual paraphrasing by experts and non-experts. |
| Outcome: | The proposed methods outperform human translation systems in a variety of translation tasks. |
Appraise Evaluation Framework for Machine Translation (C18-2)
Copied to clipboard
| Challenge: | Appraise is an open-source framework for crowd-based annotation tasks . it is used for shared tasks at the conference on machine translation and at IWSLT 2017 . |
| Approach: | They present an open-source framework for crowd-based annotation tasks . they describe the entire lifecycle of an Appraise evaluation campaign . |
| Outcome: | The proposed framework is used to run evaluation campaigns at the WMT Conference on Machine Translation and at IWSLT 2017 . it has been adopted by the translator team at Microsoft Translator for internal quality monitoring . |
Assessing Human-Parity in Machine Translation on the Segment Level (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent machine translation shared tasks have shown top-performing systems to tie or outperform human translation. |
| Approach: | They examine the outputs of top-performing systems in a recent machine translation shared task . they find that some systems outperform human translation on average . |
| Outcome: | a new method identifies segments for which human and machine perform poorly . the results show that top-performing systems outperform human translation on average . |
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies (2024.acl-long)
Copied to clipboard
| Challenge: | a decade ago a single metric, BLEU, governed progress in machine translation research. |
| Approach: | They investigate the "dynamic range" of a number of modern machine translation metrics to provide a collective understanding of differences in scores . they use a large dataset to discover deltas at which metrics achieve system-level differences that are meaningful to humans . |
| Outcome: | The proposed method is more stable than statistical p-values in regards to testset size. |