Challenge: Recent advances in natural language processing and computer vision have made it possible to translate images with text in one language into equivalent images displaying that text translated into another language.
Approach: They propose an all-encompassing framework for the task–In-Image Machine Translation (IIMT) that incorporates contextual cues from both textual and visual elements during translation.
Outcome: The proposed framework can be constructed using open-source models and requires no training, making it highly accessible and expandable.

Similar Papers

MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation (2025.coling-main)

Copied to clipboard

Challenge: Existing datasets suffer from limitations in scale, diversity, and quality, hindering the development and evaluation of IT models.
Approach: They propose a large-scale parallel corpus of multilingual image translation with over 10M image-text pairs derived from real-world data.
Outcome: The proposed model performs better in tackling challenging and complex image translation tasks in the real world.
PRIM: Towards Practical In-Image Multilingual Machine Translation (2025.emnlp-main)

Copied to clipboard

Challenge: Current research on in-image machine translation focuses on synthetic data with simple background, single font, fixed text position, and bilingual translation.
Approach: They propose an end-to-end model to handle the challenge of practical conditions in PRIM . they annotate a real-world one-line text image with complex background, fonts, diverse text positions .
Outcome: The proposed model improves translation quality and visual effect compared to other models.
In-Image Machine Translation. A Preliminary Modular Approach (2026.eacl-srw)

Copied to clipboard

Challenge: In-image machine translation is a sub-task of Image-Based Machine Translation that aims to substitute text embedded in images with its translation into another language.
Approach: They propose a simple task that renders parallel text over a plain background and a pipeline that obtains the transcript of the original image, translates it, and generates a new image similar to the original one.
Outcome: The proposed approach outperforms existing models including an end-to-end approach and is competitive with other similar approaches.
Learning Translations via Images with a Massively Multilingual Image Dataset (P18-1)

Copied to clipboard

Challenge: Existing datasets for learning translations of words are limited to a few high-resource languages and unrealistically easy settings.
Approach: They propose a large-scale multilingual corpus of images labeled with the word they represent to facilitate translation research.
Outcome: The proposed method improves on an unsupervised technique that has been limited to a few languages and unrealistic settings.
Exploring In-Image Machine Translation with Real-World Background (2025.findings-acl)

Copied to clipboard

Challenge: Existing models for IIMT focus on simplified scenarios, which is far from reality and impractical for applications in the real world.
Approach: They propose a model that separates the background and text-image from the source image and performs translation on the text- image directly.
Outcome: The proposed model improves translation quality and visual effect in complex scenarios . it separates background and text-image from source image and performs translation on the text- image directly .
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer (2021.naacl-main)

Copied to clipboard

Challenge: Current natural language processing pipelines often use transfer learning, where a model is pre-trained on a data-rich task before being fine-tuned on . this significantly limits their use given that roughly 80% of the world population does not speak English.
Approach: They introduce a multilingual variant of T5 that was pre-trained on a new Common Crawl-based dataset covering 101 languages.
Outcome: The proposed model achieves state-of-the-art on multilingual benchmarks and a simple technique to prevent accidental translation in the zero-shot setting.
An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance (2024.emnlp-main)

Copied to clipboard

Challenge: a new task is to translate images to make them culturally relevant . currently, translation systems focus on translating words and images .
Approach: They propose a task of translating images to make them culturally relevant . they build pipelines comprising state-of-the-art generative models to do the task .
Outcome: The proposed pipelines can translate only 5% of translated images for some countries and no translation is successful for others.
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown continuously improving multilingual capabilities.
Approach: They evaluate the ability of open LLMs to handle multilingual machine translation tasks using a parallel-first monolingual-second data mixing strategy.
Outcome: The proposed model outperforms state-of-the-art models and achieves competitive performance with Google Translate and GPT-4-turbo.
Translation-Enhanced Multilingual Text-to-Image Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing models for text-to-image generation are mostly based on the English language due to the lack of annotated image-caption data in other languages.
Approach: They propose to use a multilingual multi-modal encoder to bootstrap mTTI systems that can be translated into other languages.
Outcome: The proposed approach mitigates the language gap and improves on standard mTTI datasets.
Translatotron-V(ison): An End-to-End Model for In-Image Machine Translation (2024.findings-acl)

Copied to clipboard

Challenge: In-image machine translation (IIMT) aims to translate an image containing texts in source language into an image with translations in target language.
Approach: They propose an end-to-end IIMT model with four modules that translate images . they propose a two-stage training framework to assist the model in learning alignment across languages .
Outcome: The proposed model outperforms cascaded models with only 70.9% of parameters and is highly accurate.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations