Challenge: Current studies on text image translation face bottlenecks due to lack of a publicly available dataset and poor optical character recognition.
Approach: They propose a text image translation model with a multimodal codebook and an OCR dataset for Chinese-English translation.
Outcome: The proposed model can associate the image with relevant texts, providing useful supplementary information for translation.

Similar Papers

UMTIT: Unifying Recognition, Translation, and Generation for Multimodal Text Image Translation (2024.lrec-main)

Copied to clipboard

Challenge: Current Image machine translation (IMT) relies on a cascaded system that combines Optical Character Recognition (OCR) and a complex process of rendering the translated text back onto the source image.
Approach: They propose a multimodal image-text translation model that generates consistent target images . they use two image-to-text conversion steps to convert images to text to recognize source text .
Outcome: The proposed model outperforms existing methods and surpasses state-of-the-art methods in text recognition tasks.
Document Image Machine Translation with Dynamic Multi-pre-trained Models Assembling (2024.naacl-long)

Copied to clipboard

Challenge: Existing TIMT tasks focus on text-line-level images.
Approach: They propose to extend the existing TIMT task and introduce a new framework to translate a source document image to markdown-formatted target translation.
Outcome: The proposed task aims to translate a source document image with long context and complex layout structure to markdown-formatted target translation.
Translation-Enhanced Multilingual Text-to-Image Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing models for text-to-image generation are mostly based on the English language due to the lack of annotated image-caption data in other languages.
Approach: They propose to use a multilingual multi-modal encoder to bootstrap mTTI systems that can be translated into other languages.
Outcome: The proposed approach mitigates the language gap and improves on standard mTTI datasets.
Beyond Triplet: Leveraging the Most Data for Multimodal Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Multimodal machine translation (MMT) aims to improve translation quality by incorporating information from other modalities, such as vision.
Approach: They propose a framework for multimodal machine translation that utilizes large-scale non-triple data and a multimodal translation dataset.
Outcome: The proposed method can significantly improve translation performance with more non-triple data.
In-Image Machine Translation. A Preliminary Modular Approach (2026.eacl-srw)

Copied to clipboard

Challenge: In-image machine translation is a sub-task of Image-Based Machine Translation that aims to substitute text embedded in images with its translation into another language.
Approach: They propose a simple task that renders parallel text over a plain background and a pipeline that obtains the transcript of the original image, translates it, and generates a new image similar to the original one.
Outcome: The proposed approach outperforms existing models including an end-to-end approach and is competitive with other similar approaches.
MT3: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation (2026.acl-long)

Copied to clipboard

Challenge: Text Image Machine Translation (TIMT) is a critical subfield of machine translation . it requires accurate optical character recognition, robust visual-text reasoning, and high-quality translation a challenge .
Approach: They propose a multi-task optimization framework to specialize MLLMs into expert TIMT models.
Outcome: The proposed model outperforms baselines on the latest in-domain MIT-10M benchmark.
Cross-lingual Cross-modal Pretraining for Multimodal Retrieval (2021.naacl-main)

Copied to clipboard

Challenge: Recent pretrained vision-language models have achieved impressive performance on cross-modal retrieval tasks in English.
Approach: They propose a new approach to learn cross-lingual cross-modal representations for matching images and captions in multiple languages using an annotated corpus.
Outcome: The proposed model achieves impressive performance on two multimodal multilingual image caption benchmarks: Multi30k with German captions and MSCOCO with Japanese captions.
PEIT: Bridging the Modality Gap with Pre-trained Models for End-to-End Image Translation (2023.acl-long)

Copied to clipboard

Challenge: Image translation is a task that translates an image containing text in the source language to the target language.
Approach: They propose an end-to-end image translation framework that bridges the modality gap between visual inputs and textual inputs/outputs of machine translation (MT).
Outcome: The proposed framework outperforms existing models on a large-scale image translation corpus . it significantly outperformed both cascaded and strong models on the e-commerce domain .
Multilingual Multimodal Learning with Machine Translated Text (2022.findings-emnlp)

Copied to clipboard

Challenge: Currently, most vision-and-language pretraining research focuses on English tasks due to the availability of datasets.
Approach: They propose a framework for machine translating English multimodal data to improve training data . they propose two metrics to prevent models from learning from low-quality translated text .
Outcome: The proposed framework can be applied to any multimodal dataset and model.
AnyTrans: Translate AnyText in the Image with Large Scale Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in natural language processing and computer vision have made it possible to translate images with text in one language into equivalent images displaying that text translated into another language.
Approach: They propose an all-encompassing framework for the task–In-Image Machine Translation (IIMT) that incorporates contextual cues from both textual and visual elements during translation.
Outcome: The proposed framework can be constructed using open-source models and requires no training, making it highly accessible and expandable.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations