Challenge: Simultaneous interpretation is a cognitively taxing task, and even seasoned professionals benefit from real-time assistance.
Approach: They propose a simultaneous interpretation task that mimics the cognitive load of interpretation with crowdworker surrogates.
Outcome: The proposed task mimics the cognitive load of interpretation with crowdworker surrogates . the evaluation setup provides consistent results between expert and proxy participants .

Similar Papers

Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies (2026.findings-acl)

Copied to clipboard

Challenge: Simultaneous machine translation requires high-quality translations under strict real-time constraints.
Approach: They extend the action space of simultaneous machine translation with four adaptive actions . they adapt these actions in a large language model framework and construct training references .
Outcome: The proposed framework improves semantic metrics and achieves lower delay compared to reference translations and salami-based baselines.
Barriers to Effective Evaluation of Simultaneous Interpretation (2024.findings-eacl)

Copied to clipboard

Challenge: Existing studies have relied on out-of-the-box machine translation metrics to evaluate interpretation data, but they do not account for human judgments of interpretation quality.
Approach: They propose to use machine translation metrics to evaluate human interpretations to address potential barriers to disfluency, summarization, paraphrasing and segmentation.
Outcome: The proposed model achieves better correlation with human judgments than state-of-the-art metrics.
Automatic Estimation of Simultaneous Interpreter Performance (P18-2)

Copied to clipboard

Challenge: Existing methods to predict interpreter confidence and the adequacy of the interpreted message are lacking.
Approach: They propose to extend a QE pipeline to estimate interpreter performance by using five settings in three language pairs.
Outcome: The proposed method can predict interpreter confidence and adequacy over five settings in three language pairs and improves interpretation strategy and evaluation measures.
SimQA: Detecting Simultaneous MT Errors through Word-by-Word Question Answering (2022.emnlp-main)

Copied to clipboard

Challenge: a good SimulMT system will allow the downstream QA system to answer correctly as quickly as possible.
Approach: They propose a word-by-word question answering evaluation task to evaluate if models translate salient elements of a question correctly.
Outcome: a new evaluation task aims to show whether models translate salient elements of a question accurately and quickly . evaluators can reveal weaknesses in existing neural systems, hallucinating or omitting facts . human evaluation is too costly and slow to guide system development, authors say .
Exploiting Multimodal Reinforcement Learning for Simultaneous Machine Translation (2021.eacl-main)

Copied to clipboard

Challenge: Existing studies on multimodality in simultaneous machine translation have highlighted the challenges for the agent to maintain good translation quality while learning an optimal translation path.
Approach: They propose a multimodal approach to simultaneous machine translation using reinforcement learning with strategies to integrate visual and textual information in both the agent and the environment.
Outcome: The proposed multimodal approach improves translation quality while keeping latency low while providing visual cues.
Lost in Interpretation: Predicting Untranslated Terminology in Simultaneous Interpretation (N19-1)

Copied to clipboard

Challenge: Experimental results on a newly-annotated version of the NAIST Simultaneous Translation Corpus indicate the promise of our proposed method.
Approach: They propose a task of predicting which terminology simultaneous interpreters will leave untranslated using supervised sequence taggers.
Outcome: The proposed method predicts which terminology interpreters leave untranslated . it is based on an annotated version of the NAIST Simultaneous Translation Corpus .
It Is Not As Good As You Think! Evaluating Simultaneous Machine Translation on Interpretation Data (2021.emnlp-main)

Copied to clipboard

Challenge: Existing siMT systems are trained and evaluated on offline translations . however, evaluation gap remains notable, calling for constructing large-scale interpretation corpora .
Approach: They propose a translation-to-interpretation transfer method which converts offline translations into interpretation-style data.
Outcome: The proposed interpretation test set shows that SiMT models improve on translation vs interpretation data.
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair (2024.emnlp-main)

Copied to clipboard

Challenge: Existing siMT corpora are limited due to high costs and limited annotator capabilities.
Approach: They propose a method to convert ST corpora into interpretation-style corpors by fine-tuning models with Large Language Models.
Outcome: The proposed method reduces latency while achieving better quality compared to other models.
Toward Machine Interpreting: Lessons from Human Interpreting Studies (2025.emnlp-main)

Copied to clipboard

Challenge: Current speech translation systems are static and do not adapt to real-world situations in ways human interpreters do.
Approach: They propose to model human interpreting using a new language model to improve usability . they argue that there is great potential to adopt many human interpreted principles .
Outcome: The proposed models can be used to improve human interpreting and improve translation performance.
What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks? (2021.acl-long)

Copied to clipboard

Challenge: Despite the importance of datasets for natural language understanding, there has been little attention on crowdsourcing methods for collecting datasets.
Approach: They compare the effectiveness of crowdsourcing methods for boosting NLU example difficulty with training crowdworkers instead of expert judgments.
Outcome: The proposed method is ineffective for boosting NLU example difficulty, but it is not effective for training crowdworkers and qualifying workers based on expert judgments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations