Challenge: Existing studies on multimodality in simultaneous machine translation have highlighted the challenges for the agent to maintain good translation quality while learning an optimal translation path.
Approach: They propose a multimodal approach to simultaneous machine translation using reinforcement learning with strategies to integrate visual and textual information in both the agent and the environment.
Outcome: The proposed multimodal approach improves translation quality while keeping latency low while providing visual cues.

Similar Papers

Simultaneous Machine Translation with Visual Context (2020.emnlp-main)

Copied to clipboard

Challenge: Simultaneous machine translation (SiMT) aims to reproduce human interpretation, where an interpreter translates spoken utterances as they are produced.
Approach: They propose to add visual context to siMT to compensate for the missing source context . they show visual-grounded models are much better than commonly used global features .
Outcome: The proposed models reach up to 3 BLEU points improvement under low latency scenarios.
A Generative Framework for Simultaneous Machine Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches use a fixed number of source words to translate or learn dynamic policies for the number of sources by reinforcement learning.
Approach: They propose a generative framework that uses a latent variable to model read or translate actions at every time step and integrates out to consider all possible translation policies.
Outcome: The proposed framework achieves the best BLEU scores on benchmark datasets.
Simultaneous Machine Translation with Tailored Reference (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing SiMT models are trained using the same reference disregarding the varying amounts of available source information at different latency.
Approach: They propose a method that provides tailored reference for the SiMT models trained at different latency by rephrasing ground-truth to the tailored reference.
Outcome: The proposed method achieves state-of-the-art translation performance on three translation tasks.
Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies (2026.findings-acl)

Copied to clipboard

Challenge: Simultaneous machine translation requires high-quality translations under strict real-time constraints.
Approach: They extend the action space of simultaneous machine translation with four adaptive actions . they adapt these actions in a large language model framework and construct training references .
Outcome: The proposed framework improves semantic metrics and achieves lower delay compared to reference translations and salami-based baselines.
Turning Fixed to Adaptive: Integrating Post-Evaluation into Simultaneous Machine Translation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to perform adaptive and fixed translations lack evaluation before taking actions.
Approach: They propose a method to perform adaptive translation policy via post-evaluation into fixed policy . their method evaluates rationality of next action by measuring change in source content .
Outcome: The proposed method exceeds strong baselines under all latency.
Simultaneous Translation with Flexible Policy via Restricted Imitation Learning (P19-1)

Copied to clipboard

Challenge: Existing approaches to simultaneous translation have been limited and use fixed-latency policies or a complicated two-staged model.
Approach: They propose a single model that adds a “delay” token to the target vocabulary and a restricted dynamic oracle to greatly simplify training.
Outcome: The proposed model achieves better BLEU scores and lower latencies compared to fixed and RL-learned policies on Chinese -> English simultaneous translation.
A Visual Attention Grounding Neural Model for Multimodal Machine Translation (D18-1)

Copied to clipboard

Challenge: Existing approaches to multimodal machine translation do not integrate visual information into the translation process.
Approach: They propose a multimodal machine translation model that utilizes parallel visual and textual information.
Outcome: The proposed model outperforms existing methods on the Multi30K and Ambiguous COCO datasets.
Probing the Need for Visual Context in Multimodal Machine Translation (N19-1)

Copied to clipboard

Challenge: Current work on multimodal machine translation (MMT) suggests that the visual modality is either unnecessary or only marginally beneficial.
Approach: They propose to use the visual modality to combine visual and textual information to generate better translations by partially depriving models from source-side textual context.
Outcome: The proposed model can combine visual and textual information to generate better translations under limited textual context.
Prediction Improves Simultaneous Neural Machine Translation (D18-1)

Copied to clipboard

Challenge: Current systems for simultaneous machine translation use an AGENT to control an incremental encoder-decoder model.
Approach: They propose a general-purpose prediction action which predicts future words in the input stream.
Outcome: The proposed agent with prediction has better translation quality and less delay compared to an agent-based system without prediction.
Beyond Triplet: Leveraging the Most Data for Multimodal Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Multimodal machine translation (MMT) aims to improve translation quality by incorporating information from other modalities, such as vision.
Approach: They propose a framework for multimodal machine translation that utilizes large-scale non-triple data and a multimodal translation dataset.
Outcome: The proposed method can significantly improve translation performance with more non-triple data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations