Challenge: Existing synthesis methods cannot guarantee data quality.
Approach: They propose a hierarchical reward that balances translation quality and latency objectives by combining supervised fine-tuning data with supervised inputs.
Outcome: The proposed model can reuse key-value caches across both modalities and eliminate redundant feature recomputation.

Similar Papers

Learning Adaptive Segmentation Policy for End-to-End Simultaneous Translation (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to perform simultaneous speech-to-text translation ignore contextual information and suffer from low translation quality.
Approach: They propose an adaptive segmentation policy for simultaneous speech-to-text translation . it learns to segment the source streaming speech into meaningful units .
Outcome: The proposed method achieves a good accuracy-latency trade-off over state-of-the-art methods on English-German and Chinese-English.
InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model (2025.findings-acl)

Copied to clipboard

Challenge: Existing models for simultaneous speech translation assume pre-segmented speech, limiting their real-world applicability.
Approach: They propose a multi-turn dialogue task that can translate unbounded streaming speech . they construct translation trajectories and robust segments from MuST-C with multi-latency augmentation during training and develop a cache management strategy to facilitate efficient inference.
Outcome: The proposed approach reduces computation-aware latency by 0.5 to 1 second while maintaining the same translation quality compared to baselines.
SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation (2025.findings-acl)

Copied to clipboard

Challenge: SeqPO-SiMT is a new policy optimization framework for simultaneous machine translation that combines a tailored reward with a single step task.
Approach: They propose a new policy optimization framework that defines the simultaneous machine translation task as a sequential decision making problem with a tailored reward.
Outcome: The proposed framework outperforms the supervised fine-tuning model by 1.13 points while reducing the Average Lagging by 6.17 in the NEWSTEST2021 En Zh dataset.
Simultaneous Translation Policies: From Fixed to Adaptive (2020.acl-main)

Copied to clipboard

Challenge: Adaptive policies can balance translation quality and latency based on context information . previous methods on obtaining adaptive policies rely on complicated training process .
Approach: They propose to obtain adaptive policies by a simple heuristic composition of fixed policies . they propose to use a heurism to obtain policies that can outperform fixed ones .
Outcome: Experiments on Chinese -> English and German -> english show that adaptive policies outperform fixed policies by up to 4 BLEU points for the same latency.
Learning Optimal Policy for Simultaneous Machine Translation via Binary Search (2023.acl-long)

Copied to clipboard

Challenge: Simultaneous machine translation model needs a precise translation policy to achieve good latency-quality trade-offs.
Approach: They propose a method for building the optimal translation policy online via binary search by employing explicit supervision.
Outcome: Experiments on four translation tasks show that the proposed method exceeds strong baselines across all latency scenarios.
LLMs Can Achieve High-quality Simultaneous Machine Translation as Efficiently as Offline (2025.findings-acl)

Copied to clipboard

Challenge: Large language models perform well in offline machine translation when the complete source sentence is provided . however, in many real scenarios, the source tokens arrive in a streaming manner and simultaneous machine translation is required .
Approach: They propose a new paradigm that includes constructing supervised fine-tuning data for simultaneous machine translation (SiMT) to achieve SiMT, source and target tokens are rearranged into interleaved sequences, separated by special tokens according to varying latency requirements.
Outcome: The proposed approach achieves state-of-the-art performance across various SiMT benchmarks and evaluation metrics while maintaining efficient auto-regressive decoding.
Simul-MuST-C: Simultaneous Multilingual Speech Translation Corpus Using Large Language Model (2024.emnlp-main)

Copied to clipboard

Challenge: Simultaneous speech translation (SiST) begins translating before the entire source input is received.
Approach: They propose a dataset that rearranges sentences into segmented monotonic data for simultaneous speech translation using the Large Language Model.
Outcome: The proposed dataset improves quality and latency in siST translations by rearranging sentences into segmented monotonic data.
Simpler and Faster Learning of Adaptive Policies for Simultaneous Translation (D19-1)

Copied to clipboard

Challenge: Recent work on simultaneous translation is difficult because of its latency and quality.
Approach: They propose a supervised-learning framework to learn adaptive policies from parallel text sequences . they use a model that predicts when a target word is read or WRITE if context provides enough information .
Outcome: Experiments on German=>English show that the proposed method can learn flexible policies with better BLEU scores and similar latencies compared to previous work.
Simultaneous Translation with Flexible Policy via Restricted Imitation Learning (P19-1)

Copied to clipboard

Challenge: Existing approaches to simultaneous translation have been limited and use fixed-latency policies or a complicated two-staged model.
Approach: They propose a single model that adds a “delay” token to the target vocabulary and a restricted dynamic oracle to greatly simplify training.
Outcome: The proposed model achieves better BLEU scores and lower latencies compared to fixed and RL-learned policies on Chinese -> English simultaneous translation.
SegAugment: Maximizing the Utility of Speech Translation Data with Segmentation-based Augmentations (2023.findings-emnlp)

Copied to clipboard

Challenge: End-to-end Speech Translation models are limited by a data bottleneck . end-to end models can address several shortcomings of cascaded models .
Approach: They propose a data augmentation strategy to augment sentence-level datasets by using an Audio Segmentation system to re-segment the speech of each document with different length constraints.
Outcome: The proposed method achieves state-of-the-art results in MuST-C and in mTEDx.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations