Papers with VMT

4 papers
SHIFT: Selected Helpful Informative Frame for Video-guided Machine Translation (2025.emnlp-main)

Copied to clipboard

Challenge: Video-guided machine translation (VMT) aims to improve translation quality by integrating contextual information from paired short video clips.
Approach: They propose a plug-and-play framework for video-guided machine translation with multimodal large language models.
Outcome: The proposed framework improves performance of MLLMs while reducing computational cost.
DART: Disambiguation-Aware Reasoning for Video-guided Machine Translation (2026.acl-long)

Copied to clipboard

Challenge: Video-guided Machine Translation (VMT) uses short video clips to enhance translation quality, but many samples are text-sufficient.
Approach: They propose a framework that integrates multimodal large language models’ multimodal reasoning into video-guided machine translation by using a pipeline for constructing training data based on multimodal relevance to translation.
Outcome: The proposed framework improves multimodal information utilization in video-guided machine translation, yielding gains in translation quality and computational efficiency.
TriFine: A Large-Scale Dataset of Vision-Audio-Subtitle for Tri-Modal Machine Translation and Benchmark with Fine-Grained Annotated Tags (2025.coling-main)

Copied to clipboard

Challenge: Existing video-guided machine translation approaches use coarse-grained visual information, resulting in information redundancy and high computational overhead.
Approach: They propose a fine-grained approach to video-guided machine translation using visual information . they use a large-scale dataset with annotated multimodal fine-grain tags .
Outcome: The proposed approach achieves superior performance with lower computational overhead compared to coarse-grained methods and text-only models.
The Effects of Pretraining in Video-Guided Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to improve VMT models integrate text and video modalities.
Approach: They propose an approach that improves the performance of VMT models by using a new dataset which contains transcribed audio descriptions of movies.
Outcome: The proposed model improves on the MAD (Movie Audio Descriptions) dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations