Papers with MMT

45 papers
3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have shown that visual information in existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities.
Approach: They propose to use 3AM to create an ambiguity-aware multimodal machine translation dataset.
Outcome: The proposed dataset includes more ambiguity and a greater variety of captions and images than other MMT datasets.
Choosing What to Mask: More Informed Masking for Multimodal Machine Translation (2023.acl-srw)

Copied to clipboard

Challenge: Pre-trained language models have achieved remarkable results on several NLP tasks.
Approach: They propose three new masking strategies for cross-lingual visual pre-training that focus on learning different linguistic patterns.
Outcome: The proposed methods outperform the baseline model and achieve state-of-the-art accuracy on the Portuguese-English MMT task.
Language-Aware Multilingual Machine Translation with Self-Supervised Learning (2023.findings-eacl)

Copied to clipboard

Challenge: Multilingual machine translation (MMT) is a challenging multitask optimization problem because of lack of a framework to learn language-specific parameters.
Approach: They propose a self-supervised learning task that denies monolingual data to MMT . they then propose 'intra-distillation' task that co-trains with MMT task .
Outcome: The proposed approach outperforms three state-of-the-art methods on 8-language and 15-language benchmarks.
Towards Zero-Shot Multimodal Machine Translation (2025.findings-naacl)

Copied to clipboard

Challenge: Current multimodal machine translation systems rely on fully supervised data, which is costly to collect and prevents extension of MMT to language pairs with no such data.
Approach: They propose a method to bypass the need for fully supervised data to train MMT systems . they adapt a strong text-only machine translation model to a visually conditioned language model and a divergence test set to evaluate how well models use images to disambiguate translations.
Outcome: The proposed method can generalize to languages with no fully supervised training data.
Detect, Disambiguate, and Translate: On-Demand Visual Reasoning for Multimodal Machine Translation with Large Vision-Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Multimodal machine translation (MMT) aims to leverage additional modalities beyond text . current MMT systems rely heavily on monolingual English captioning data .
Approach: They propose a reasoning-based framework to leverage large-scale vision-language models for MMT . they propose Detect, Disambiguate, and Translate framework to detect ambiguity in input sentence .
Outcome: The proposed framework outperforms state-of-the-art models in disambiguation accuracy and translation quality.
Parameter-Efficient Neural Reranking for Cross-Lingual and Multilingual Retrieval (2022.coling-1)

Copied to clipboard

Challenge: State-of-the-art neural rankers are notoriously data-hungry and rarely used in multilingual and cross-lingual retrieval settings.
Approach: They propose to use Sparse Fine-Tuning Masks and Adapters to transfer rankers trained on English data to other languages and cross-lingual setups by means of multilingual encoders.
Outcome: The proposed methods outperform standard zero-shot transfer with full MMT fine-tuning while being more modular and reducing training times.
Entity-level Cross-modal Learning Improves Multi-modal Machine Translation (2021.findings-emnlp)

Copied to clipboard

Challenge: Multi-modal machine translation aims at improving translation performance by incorporating visual information.
Approach: They propose an explicit entity-level cross-modal learning approach that aims to augment the entity representation by combining a translation task and a reconstruction task.
Outcome: The proposed approach achieves comparable or even better performance than state-of-the-art models.
Multilingual and Multimodal Learning for Brazilian Portuguese (2022.lrec-1)

Copied to clipboard

Challenge: Existing models that learn multimodal and multilingual representations perform better in many natural language tasks.
Approach: They use a multimodal and multilingual corpus to test its generalization ability for other languages . they achieve a BLEU score of 51.8 and a METEOR score of 78.0 on the test set .
Outcome: The proposed model outperforms the existing model on a Portuguese-English multimodal translation task.
Disentangling the Roles of Target-side Transfer and Regularization in Multilingual Machine Translation (2024.eacl-long)

Copied to clipboard

Challenge: Multilingual Machine Translation (MMT) benefits from knowledge transfer across different language pairs, but performance differences between one-to-many and many-to-1 translation are negligible.
Approach: They conduct a large-scale study that varies the auxiliary target-side languages along two dimensions to show the dynamic impact of knowledge transfer on the main language pairs.
Outcome: The proposed model can translate between multiple languages with minimal positive transfer ability.
Efficiently Upgrading Multilingual Machine Translation Models to Support More Languages (2023.eacl-main)

Copied to clipboard

Challenge: Existing multilingual machine translation models need to be upgraded as data becomes available in more languages.
Approach: They propose three techniques that speed up the effective learning of new languages and alleviate catastrophic forgetting .
Outcome: The proposed techniques exceed the performance of a same-sized baseline model with 30% computation and recover the performance a larger model trained from scratch with over 50% reduction in computation.
Whisper-UT: A Unified Translation Framework for Speech and Text (2025.emnlp-main)

Copied to clipboard

Challenge: Encoder-decoder models have achieved remarkable success in speech and text tasks, but efficiently adapting them to diverse uni/multimodal scenarios remains a challenge.
Approach: They propose a framework that leverages lightweight adapters to enable seamless adaptation across tasks.
Outcome: The proposed framework improves speech translation performance through a 2-stage decoding strategy without requiring 3-way parallel data.
Distill The Image to Nowhere: Inversion Knowledge Distillation for Multimodal Machine Translation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies on multimodal machine translation (MMT) have focused on the fusion and alignment of images and texts to improve MMT.
Approach: They propose an image-free inference framework that supports image-based inference via an inversion knowledge distillation scheme.
Outcome: The proposed framework is the first to rival or surpass image-must frameworks on the multimodal translation benchmark.
Beyond Triplet: Leveraging the Most Data for Multimodal Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Multimodal machine translation (MMT) aims to improve translation quality by incorporating information from other modalities, such as vision.
Approach: They propose a framework for multimodal machine translation that utilizes large-scale non-triple data and a multimodal translation dataset.
Outcome: The proposed method can significantly improve translation performance with more non-triple data.
Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models require associated image with input sentence, which is difficult to satisfy at inference.
Approach: They propose to use synthetic and authentic images to generate translations using text-to-image generation models.
Outcome: The proposed model achieves state-of-the-art performance on En-De and En-Fr datasets while remaining independent of authentic images during inference.
Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis (2024.findings-naacl)

Copied to clipboard

Challenge: Existing studies show that large language models (LLMs) can handle multilingual machine translation (MMT) However, the multilingual translation ability of LLMs remains under-explored.
Approach: They evaluate eight popular LLMs including ChatGPT and GPT-4 to determine their performance in multilingual machine translation.
Outcome: The proposed model can generate moderate translation even on zero-resource languages and cross-lingual exemplars can provide better task guidance for low-resourced translation than exemplar in the same language pairs.
LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine Translation (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances struggle to train a separate model for each language pair, which is costly and unaffordable when the number of languages increases in the real world.
Approach: They propose to train different MMT models to support translations between different languages.
Outcome: The proposed model is able to handle the above issues by providing a shared semantic space for multiple languages.
Video-Helpful Multimodal Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal machine translation datasets contain images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity.
Approach: They propose an MMT dataset that contains ambiguous subtitles and a video-helpful evaluation set.
Outcome: The proposed model performs significantly better than existing models on ambiguous subtitles dataset . it is based on a training set and video-helpful evaluation set .
Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation (2023.acl-long)

Copied to clipboard

Challenge: Recent work in multimodal machine translation (MT) has shown that ambiguity can be resolved using accompanying context such as images.
Approach: They propose a multimodal machine translation approach based on a strong text-only MT model and a novel guided self-attention mechanism to train it.
Outcome: The proposed model outperforms existing models on EnglishFrench, EnglishGerman and EnglishCzech benchmarks and is freely available.
Probing Multi-modal Machine Translation with Pre-trained Language Model (2021.findings-acl)

Copied to clipboard

Challenge: Multi-modal machine translation (MMT) aimed at using images to help disambiguate the target during translation but recent studies showed that visual features are either negligible or incremental.
Approach: They propose to incorporate a visual language model on the source side to improve multi-modal translation quality significantly.
Outcome: The proposed model improves the translation quality significantly on the multi-modal dataset.
When Does Monolingual Data Help Multilingual Translation: The Role of Domain and Model Scale (2024.naacl-long)

Copied to clipboard

Challenge: Multilingual machine translation (MMT) is a key tool for improving translation in low-resource languages.
Approach: They examine how denoising autoencoding and backtranslation impact multilingual machine translation under different data conditions and model scales.
Outcome: The proposed method improves translation efficiency in low-resource languages by using denoising autoencoding (DAE) and backtranslation (BT) .
Alternative Input Signals Ease Transfer in Multilingual Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Recent work in multilingual machine translation (MMT) has focused on the potential of positive transfer between languages.
Approach: They propose to augment training data with alternative signals that unify different writing systems, such as phonetic, romanized, and transliterated input.
Outcome: The proposed model outperforms strong ensemble baselines on Indic and Turkic languages by 1.3 BLEU points on both languages.
From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformers (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that multilingual transformers are less effective in resource-lean scenarios and for distant languages.
Approach: They propose to use massively multilingual transformers to pretrain languages . they show that MMTs are less effective in resource-lean scenarios and distant languages if they are pre-trained via language modeling .
Outcome: The proposed model is less effective in resource-lean scenarios and for distant languages than cross-lingual word embeddings.
Neural Machine Translation with Phrase-Level Universal Visual Representations (2022.acl-long)

Copied to clipboard

Challenge: Existing multimodal machine translation methods require paired input of source sentence and image, which makes them suffer from shortage of sentence-image pairs.
Approach: They propose a phrase-level retrieval-based method to get visual information from existing sentence-image data sets.
Outcome: The proposed method significantly outperforms strong baselines on multiple MMT datasets, especially when the textual context is limited.
Multimodal Transformer for Multimodal Machine Translation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to incorporate information from other modality, usually static images, are not considered relative to multimodal machine translation.
Approach: They propose a multimodal self-attention method which learns the representation of images based on the text, which avoids encoding irrelevant information in images.
Outcome: The proposed model outperforms previous studies and competitive baselines in terms of various metrics.
Probing the Need for Visual Context in Multimodal Machine Translation (N19-1)

Copied to clipboard

Challenge: Current work on multimodal machine translation (MMT) suggests that the visual modality is either unnecessary or only marginally beneficial.
Approach: They propose to use the visual modality to combine visual and textual information to generate better translations by partially depriving models from source-side textual context.
Outcome: The proposed model can combine visual and textual information to generate better translations under limited textual context.
On Vision Features in Multimodal Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Recent work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is given to the quality of vision models.
Approach: They develop a selective attention model to study the patch-level contribution of an image in multimodal machine translation.
Outcome: The proposed model is able to learn translation from the visual modality on probing tasks and is compared with existing models.
Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic Perspective (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on extracting multi-granularity visual features for integration or designing model architectures for better message passing across various modalities.
Approach: They propose to decompose the informative visual signals into two parts: source-specific information and target-specific info.
Outcome: The proposed method can enhance the visual awareness of MMT models against strong baselines.
Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: Recent studies report improvements when equipping models with multimodal information, but it remains unclear whether such improvements actually come from the multimodal part.
Approach: They propose to extend conventional text-only translation models with multimodal information by extending them with visual input.
Outcome: The proposed models replicate similar gains as recently developed multimodal-integrated systems achieved, but learn to ignore multimodal information.
Multimodal Machine Translation with Text-Image In-depth Questioning (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal machine translation (MMT) models focus on intermodal interactions, but focus on simple interactions between nouns and entities in image, overlooking global semantic alignment.
Approach: They propose a Text-Image In-depth Questioning method to deepen interactions and optimize translations by utilizing visual data to capture global semantic alignment.
Outcome: The proposed method achieves state-of-the-art results on five translation directions of Multi30K and AmbigCaps, with +2.35 BLEU on the challenging MSCOCO benchmark.
Distilling Efficient Language-Specific Models for Cross-Lingual Transfer (2023.findings-acl)

Copied to clipboard

Challenge: Massively multilingual Transformers (MMTs) are widely used for cross-lingual transfer learning.
Approach: They propose to extract compressed, language-specific models from MMTs which retain the capacity of the original MMT for cross-lingual transfer.
Outcome: The proposed model outperforms models trained from scratch in zero-shot cross-lingual transfer across benchmarks.
VQA-Augmented Machine Translation with Cross-Modal Contrastive Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multimodal machine translation methods often extract visual features using pre-trained models while learning text features from scratch, leading to representation imbalance.
Approach: They propose a cross-modal VQA-augmented multimodal machine translation method . it aligns image-source text pairs and image-question text pairs through dual-text contrastive learning .
Outcome: The proposed method outperforms state-of-the-art methods on multiple evaluation metrics.
Imagination and Contemplation: A Balanced Framework for Semantic-Augmented Multimodal Machine Translation (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal Machine Translation (MMT) is effective in resolving linguistic ambiguities, but visual information often introduces redundancy or noise, potentially impairing translation quality.
Approach: They propose a semantic-augmented framework that integrates "Imagination" and "Contemplation" they first generate synthetic images from source text and align them with authentic images via an optimal transport loss .
Outcome: The proposed framework outperforms baselines on translation datasets with visually ambiguous or weakly correlated content.
Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup (2024.acl-long)

Copied to clipboard

Challenge: Multimodal machine translation (MMT) aims to improve the performance of machine translation with the help of visual information.
Approach: They propose a multimodal machine translation mixup method that integrates visual information into conventional text-only neural machine translation systems.
Outcome: The proposed method outperforms existing models on a multi-directional dataset with fewer parameters and achieves new state-of-the-art performance.
Latent Variable Model for Multi-modal Translation (P19-1)

Copied to clipboard

Challenge: Libovick and Helcl (2017) show improvements due to imposing a constraint on the KL term to promote models with non-negligible mutual information between inputs and latent variable and training on additional target-language image descriptions.
Approach: They propose to model interaction between visual and textual features for multi-modal neural machine translation (MMT) using a latent variable model.
Outcome: The proposed model improves over baselines including a multi-task learning approach and a conditional variational auto-encoder approach.
Multi-source Meta Transfer for Low Resource Multiple-Choice Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Existing MCQA datasets are small in size, which increases difficulty of model learning and generalization.
Approach: They propose a multi-source meta transfer framework for low-resource multiple-choice question answering . they extend meta learning by incorporating multiple training sources to learn a generalized feature representation across domains .
Outcome: The proposed framework is independent of backbone language models and can bridge the distribution gap between training sources and target.
Vision Matters When It Should: Sanity Checking Multimodal Machine Translation Models (2021.emnlp-main)

Copied to clipboard

Challenge: Multimodal machine translation models outperform text-only models when visual context is available, but recent studies have shown that the performance of MMT models is only marginally impacted when the associated image is replaced with an unrelated image or noise.
Approach: They propose to use visual data to highlight the importance of visual inputs in MMT models to enhance their leverage.
Outcome: The proposed models outperform text-only models when visual context is available, but the results show that the visual context might not be exploited by the models at all.
CCEval: A Representative Evaluation Benchmark for the Chinese-centric Multilingual Machine Translation (2023.findings-emnlp)

Copied to clipboard

Challenge: Multilingual machine translation (MMT) has gained more importance due to international business development and cross-cultural exchanges.
Approach: They propose to use Chinese-centric MMT evaluation dataset to build an impartial and representative evaluation benchmark.
Outcome: The proposed dataset covers more diverse linguistic features than other benchmarks and is highly representative and humancorrelated.
Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation (2022.lrec-1)

Copied to clipboard

Challenge: Hausa is considered a low resource language in natural language processing due to lack of resources.
Approach: They propose a dataset that contains the description of an image in Hausa and its equivalent in English.
Outcome: The Hausa Visual Genome is the first dataset of its kind . it can be used for Hausa-English machine translation, multi-modal research, image description .
VISA: An Ambiguous Subtitles Dataset for Visual Scene-aware Machine Translation (2022.lrec-1)

Copied to clipboard

Challenge: Existing multimodal machine translation datasets contain images and video captions or general subtitles which rarely contain linguistic ambiguity.
Approach: They propose a dataset that consists of Japanese-English parallel sentence pairs and corresponding video clips.
Outcome: The proposed dataset is challenging for the latest MMT system and can facilitate MMT research.
Unsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting (2020.acl-main)

Copied to clipboard

Challenge: Unsupervised machine translation (MT) has recently achieved impressive results with monolingual corpora.
Approach: They propose to utilize visual content for disambiguation and promoting latent space alignment in unsupervised machine translation by using multimodal back-translation and pseudo visual pivoting.
Outcome: The proposed model improves over state-of-the-art methods and generalizes well when images are not available at the testing time.
LAMBDA: Large Language Model-Based Data Augmentation for Multi-Modal Machine Translation (2024.findings-emnlp)

Copied to clipboard

Challenge: Multi-modal machine translation methods are underperforming compared to pre-trained models due to lack of triplet training data.
Approach: They propose a multi-modal machine translation method that integrates images and visual modality to enhance language understanding.
Outcome: The proposed method can enrich the original samples and expand the dataset without requiring external images and text.
How Far can 100 Samples Go? Unlocking Zero-Shot Translation with Tiny Multi-Parallel Data (2024.findings-acl)

Copied to clipboard

Challenge: a common solution to zero-shot translation is to add as many related translation directions as possible to the training corpus.
Approach: They show that a small amount of multi-parallel data can achieve significant zero-shot improvements . they say that the resulting non-English performance is close to the complete translation upper bound .
Outcome: The proposed model achieves +21.7 ChrF++ non-English translation improvements on EC30 dataset . the resulting non- English performance exceeds M2M100 by an average of 5.9 ChrF+ .
Asymmetric Conflict and Synergy in Post-training for LLM-based Multilingual Machine Translation (2025.findings-acl)

Copied to clipboard

Challenge: Existing work in LLM-based MMT typically mitigates the Curse of Multilinguality . asymmetric phenomenon in linguistic conflicts and synergy varies in different translation directions .
Approach: They propose a direction-aware training approach to address asymmetry in linguistic conflicts and synergy . they propose X-ALMA-13B-Pretrain with multilingual pre-training to achieve comparable performance .
Outcome: The proposed method achieves comparable performance to X-ALMA-13B-Pretrain (only SFT) with fewer pretraining tokens and 17B parameters.
Incorporating Probing Signals into Multimodal Machine Translation via Visual Question-Answering Pairs (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that multimodal machine translation systems exhibit decreased sensitivity to visual information when text inputs are complete.
Approach: They propose to generate parallel VQA style pairs from source text to foster more robust cross-modal interaction.
Outcome: The proposed approach generates parallel VQA style pairs from the source text, fostering more robust cross-modal interaction.
NiuTrans.LMT: Toward Inclusive and Scalable Multilingual Machine Translation with LLMs (2026.acl-long)

Copied to clipboard

Challenge: Large language models have significantly advanced Multilingual Machine Translation (MMT) yet scaling to many languages while maintaining robust performance across directions remains challenging.
Approach: They propose a strategy to reduce the number of translations in one direction . they propose auxiliary parallel sentences to promote cross-lingual transfer .
Outcome: The proposed model performs on par with or better than substantially larger baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations