Challenge: Spoken Language Translation (SLT) evaluation of machine translation (MT) quality has been investigated for decades.
Approach: They propose an open-source tool for assessing machine translation (MT) quality based on time-stamped transcripts and reference translations.
Outcome: The proposed evaluation tool is based on time-stamped transcripts and reference translations into a target language.

Similar Papers

Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, traditional evaluation methods struggle to detect subtle translation errors.
Approach: They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation.
Outcome: The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations.
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation (2023.emnlp-demo)

Copied to clipboard

Challenge: a framework to evaluate low-latency speech translations is currently only limited to specific aspects and is not able to compare different approaches.
Approach: They propose a framework to perform and evaluate low-latency speech translation in realistic conditions.
Outcome: The proposed framework evaluates various aspects of low-latency speech translation under realistic conditions.
Informative Manual Evaluation of Machine Translation Output (2020.coling-main)

Copied to clipboard

Challenge: a new method for manual evaluation of machine translation output is proposed . evaluators mark problematic parts of the translated text, not just overall scores .
Approach: They propose a method for manual evaluation of machine translation output based on marking actual issues in the translated text.
Outcome: The proposed method can be applied on any genre/domain and language pair . it can be guided by various types of quality criteria and can be used for other types of generated text.
SLURP: A Spoken Language Understanding Resource Package (2020.emnlp-main)

Copied to clipboard

Challenge: Publicly available datasets for Spoken Language Understanding (SLU) are limited.
Approach: They propose a publicly available SLU resource package that includes a multi-domain dataset in English spanning 18 domains.
Outcome: The proposed dataset is bigger and more diverse than existing datasets.
Multi-Dimensional Machine Translation Evaluation: Model Evaluation and Resource for Korean (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on MT evaluation characterize quality of output with a single number . a recent advancement in MT technologies has enabled higher-quality, more nuanced translations .
Approach: They propose a 1200-sentence MQM evaluation benchmark for English-Korean and a reference-free QE setup to evaluate the quality of the translations.
Outcome: The proposed model outperforms the existing model in style and accuracy.
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation (2024.lrec-main)

Copied to clipboard

Challenge: a meta-analysis of human evaluation for speech translation has not been conducted . noisy data and segmentation mismatches are challenges for automatic metrics .
Approach: They propose an evaluation strategy based on automatic resegmentation and direct assessment with segment context.
Outcome: The proposed evaluation strategy is robust and scores well-correlated with other types of human judgements.
SpeechAlign: A Framework for Speech Translation Alignment Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Speech-to-Speech and Speech- to-Text translation are currently dynamic areas of research.
Approach: They propose a framework to evaluate source-target alignment in speech models . they introduce a speech gold alignment dataset and introduce two new metrics .
Outcome: The proposed framework evaluates source-target alignment quality within speech models.
MTLens: Machine Translation Output Debugging (2022.lrec-1)

Copied to clipboard

Challenge: a demo demonstrates a system for quantitatively evaluating MT systems in isolation or multiple MT models collectively . performance of machine translation systems varies significantly with inputs of diverging features, such as genres, genres and surface properties.
Approach: They propose a benchmarking interface that quantitatively evaluates MT systems in isolation or collectively . the interface can be extended to include additional filters such as lexical, morphological, and syntactic features.
Outcome: The proposed system quantitatively evaluates MT systems on multiple domains and evaluation metrics.
A Benchmark for Translations Across Styles and Language Variants (2025.findings-emnlp)

Copied to clipboard

Challenge: lack of comprehensive evaluation benchmarks has hindered progress in this field . lack of evaluation benchmarking has hinder MT's ability to generate accurate outputs .
Approach: They evaluate translations across semantic preservation, cultural and regional specificity, expression style, and fluency at both the word and sentence levels.
Outcome: The proposed evaluation framework is validated on translations of state-of-the-art large language models .
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages (2026.acl-long)

Copied to clipboard

Challenge: Existing metrics have been developed and validated for English and other languages . this narrow focus leaves Indian languages largely overlooked, casting doubt on universality of current evaluation practices.
Approach: They propose a large-scale benchmark that compares 26 automatic metrics with human judgments across six major Indian languages.
Outcome: ITEM evaluates alignment of 26 automatic metrics with human judgments across six languages . authors: outliers exert significant impact on metric-human agreement, improve fidelity . they say the results offer critical guidance for advancing metric design and evaluation in Indian languages - a global market for machine translation and text summarization systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations