Challenge: Multi-modal machine translation aims at improving translation performance by incorporating visual information.
Approach: They propose an explicit entity-level cross-modal learning approach that aims to augment the entity representation by combining a translation task and a reconstruction task.
Outcome: The proposed approach achieves comparable or even better performance than state-of-the-art models.

Similar Papers

Probing Multi-modal Machine Translation with Pre-trained Language Model (2021.findings-acl)

Copied to clipboard

Challenge: Multi-modal machine translation (MMT) aimed at using images to help disambiguate the target during translation but recent studies showed that visual features are either negligible or incremental.
Approach: They propose to incorporate a visual language model on the source side to improve multi-modal translation quality significantly.
Outcome: The proposed model improves the translation quality significantly on the multi-modal dataset.
Multi-Modal Entities Matter: Benchmarking Multi-Modal Entity Alignment (2025.coling-main)

Copied to clipboard

Challenge: Existing MMEA datasets consider multi-modal data as attributes of textual entities, neglecting correlations between the multi-modal data.
Approach: They propose a multi-modal entity alignment dataset that models multi-dimensional data as textual entities in the MMKG.
Outcome: The proposed dataset can learn the structural information of entities by considering both intra-modal and cross-modal relations and infer the similarity of different types of entity pairs.
VQA-Augmented Machine Translation with Cross-Modal Contrastive Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multimodal machine translation methods often extract visual features using pre-trained models while learning text features from scratch, leading to representation imbalance.
Approach: They propose a cross-modal VQA-augmented multimodal machine translation method . it aligns image-source text pairs and image-question text pairs through dual-text contrastive learning .
Outcome: The proposed method outperforms state-of-the-art methods on multiple evaluation metrics.
Probing the Need for Visual Context in Multimodal Machine Translation (N19-1)

Copied to clipboard

Challenge: Current work on multimodal machine translation (MMT) suggests that the visual modality is either unnecessary or only marginally beneficial.
Approach: They propose to use the visual modality to combine visual and textual information to generate better translations by partially depriving models from source-side textual context.
Outcome: The proposed model can combine visual and textual information to generate better translations under limited textual context.
On Vision Features in Multimodal Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Recent work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is given to the quality of vision models.
Approach: They develop a selective attention model to study the patch-level contribution of an image in multimodal machine translation.
Outcome: The proposed model is able to learn translation from the visual modality on probing tasks and is compared with existing models.
Latent Variable Model for Multi-modal Translation (P19-1)

Copied to clipboard

Challenge: Libovick and Helcl (2017) show improvements due to imposing a constraint on the KL term to promote models with non-negligible mutual information between inputs and latent variable and training on additional target-language image descriptions.
Approach: They propose to model interaction between visual and textual features for multi-modal neural machine translation (MMT) using a latent variable model.
Outcome: The proposed model improves over baselines including a multi-task learning approach and a conditional variational auto-encoder approach.
LAMBDA: Large Language Model-Based Data Augmentation for Multi-Modal Machine Translation (2024.findings-emnlp)

Copied to clipboard

Challenge: Multi-modal machine translation methods are underperforming compared to pre-trained models due to lack of triplet training data.
Approach: They propose a multi-modal machine translation method that integrates images and visual modality to enhance language understanding.
Outcome: The proposed method can enrich the original samples and expand the dataset without requiring external images and text.
Beyond Triplet: Leveraging the Most Data for Multimodal Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Multimodal machine translation (MMT) aims to improve translation quality by incorporating information from other modalities, such as vision.
Approach: They propose a framework for multimodal machine translation that utilizes large-scale non-triple data and a multimodal translation dataset.
Outcome: The proposed method can significantly improve translation performance with more non-triple data.
Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic Perspective (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on extracting multi-granularity visual features for integration or designing model architectures for better message passing across various modalities.
Approach: They propose to decompose the informative visual signals into two parts: source-specific information and target-specific info.
Outcome: The proposed method can enhance the visual awareness of MMT models against strong baselines.
Multi-modal Contrastive Representation Learning for Entity Alignment (2022.coling-1)

Copied to clipboard

Challenge: Existing studies focus on how to utilize information from different modalities, but it is not trivial to leverage multi-modal knowledge in entity alignment because of the modality heterogeneity.
Approach: They propose a Multi-modal Contrastive Learning based Entity Alignment model which learns multiple individual representations from multiple modalities and performs contrastive learning to jointly model inter-modal and inter-modal interactions.
Outcome: The proposed model outperforms state-of-the-art models on public datasets under both supervised and unsupervised conditions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations