Papers with VMT
SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Video-guided machine translation (VMT) aims to improve translation quality by integrating contextual information from paired short video clips. |
| Approach: | They propose a plug-and-play framework for video-guided machine translation with multimodal large language models. |
| Outcome: | The proposed framework improves performance of MLLMs while reducing computational cost. |
DART: Disambiguation-Aware Reasoning for Video-guided Machine Translation (2026.acl-long)
Copied to clipboard
| Challenge: | Video-guided Machine Translation (VMT) uses short video clips to enhance translation quality, but many samples are text-sufficient. |
| Approach: | They propose a framework that integrates multimodal large language models’ multimodal reasoning into video-guided machine translation by using a pipeline for constructing training data based on multimodal relevance to translation. |
| Outcome: | The proposed framework improves multimodal information utilization in video-guided machine translation, yielding gains in translation quality and computational efficiency. |
TriFine: A Large-Scale Dataset of Vision-Audio-Subtitle for Tri-Modal Machine Translation and Benchmark with Fine-Grained Annotated Tags (2025.coling-main)
Copied to clipboard
| Challenge: | Existing video-guided machine translation approaches use coarse-grained visual information, resulting in information redundancy and high computational overhead. |
| Approach: | They propose a fine-grained approach to video-guided machine translation using visual information . they use a large-scale dataset with annotated multimodal fine-grain tags . |
| Outcome: | The proposed approach achieves superior performance with lower computational overhead compared to coarse-grained methods and text-only models. |
The Effects of Pretraining in Video-Guided Machine Translation (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to improve VMT models integrate text and video modalities. |
| Approach: | They propose an approach that improves the performance of VMT models by using a new dataset which contains transcribed audio descriptions of movies. |
| Outcome: | The proposed model improves on the MAD (Movie Audio Descriptions) dataset. |