Challenge: lack of large-scale open Japanese image-text pairs poses a significant barrier to the development of vision-language models.
Approach: They construct large-scale Japanese image-text pairs using machine translation and pre-trained CLIP models on a Japanese dataset.
Outcome: The results show that pre-trained models achieve competitive average scores on Japanese culture tasks compared to models of similar size.

Similar Papers

Harnessing PDF Data for Improving Japanese Large Multimodal Models (2025.findings-acl)

Copied to clipboard

Challenge: Large Multimodal Models (LMMs) have demonstrated strong performance in English, but their effectiveness in Japanese remains limited due to the lack of high-quality training data.
Approach: They propose a pipeline that leverages pretrained models to extract image-text pairs from PDFs . they use layout analysis, OCR, and vision-language pairing to enrich the training data .
Outcome: The proposed pipeline extracts image-text pairs from Japanese PDFs, eliminating manual annotations.
RWKV-CLIP: A Robust Vision-Language Representation Learner (2024.emnlp-main)

Copied to clipboard

Challenge: Using large image-text datasets, large-scale image-data sets have been used for visionlanguage pre-training.
Approach: They propose a framework that leverages Large Language Models to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags.
Outcome: The proposed framework can combine and refine information from web-based image-text pairs, synthetic captions, and detection tags.
Cross-lingual and Multilingual CLIP (2022.lrec-1)

Copied to clipboard

Challenge: OpenAI released CLIP, a model that relates the textual and visual domains with unprecedented accuracy.
Approach: They propose to use cross-lingual teacher learning to re-train an English textual encoder using a large dataset of images and captions.
Outcome: The proposed method outperforms baselines on multilingual image-text retrieval while retaining low cost.
Delving into the Openness of CLIP (2023.findings-acl)

Copied to clipboard

Challenge: Contrastive Language-Image Pre-training (CLIP) allows for open-vocabulary visual recognition, where the model can recognize images from an open class set in a zero-shot manner.
Approach: They propose to use image classification as an image-to-text matching task instead of discrete category IDs to achieve open-vocabulary visual recognition.
Outcome: The proposed model can recognize images from an open vocabulary in a zero-shot manner, but its performance deteriorates as the vocabulary expands.
AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to build a strong multilingual multimodal representation model are lacking in good-quality text-image pairs.
Approach: They propose a method to build a strong multilingual multimodal representation model using English text-image pairs instead of a model from scratch.
Outcome: The proposed model outperforms the original CLIP model on multilingual multimodal benchmarks.
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing large language models (LLMs) focus on general domains, with fewer advancements in Japanese biomedical LLMs.
Approach: They propose a benchmark for Japanese large language models with eight LLMs across four categories and 20 Japanese biomedical datasets for comparison.
Outcome: The proposed benchmark includes eight LLMs across four categories and 20 Japanese biomedical datasets across five tasks.
Scaling Data-Constrained Language Models with Synthetic Data (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) improve with more training data, but practical limitations on data collection constrain further scaling.
Approach: They compare three strategies to generate Japanese text, repeat the limited Japanese Web text, and use English Web text to fill the data shortfall.
Outcome: The proposed model outperforms baselines and achieves the performance achieved when the entire token budget is filled with additional organic Japanese Web text.
Learning Translations via Images with a Massively Multilingual Image Dataset (P18-1)

Copied to clipboard

Challenge: Existing datasets for learning translations of words are limited to a few high-resource languages and unrealistically easy settings.
Approach: They propose a large-scale multilingual corpus of images labeled with the word they represent to facilitate translation research.
Outcome: The proposed method improves on an unsupervised technique that has been limited to a few languages and unrealistic settings.
Release of Pre-Trained Models for the Japanese Language (2024.lrec-main)

Copied to clipboard

Challenge: democratization of AI aims to create a world where everyone can use AI . pre-trained models with high performance in Japanese are lagging in non-English-speaking communities .
Approach: et al. released large-scale pre-trained models trained on large-data to improve access to AI . authors say the models are more accurate and more accurate than those trained in the English language . e-mail protected: email protected.
Outcome: a new study shows that pre-trained models specialized for Japanese can achieve high performance in Japanese tasks.
Context-Informed Machine Translation of Manga using Multimodal Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Automated manga translation is a promising potential solution, but it is underdeveloped due to the need to incorporate visual elements into the translation process to resolve ambiguities.
Approach: They propose a method that leverages the vision component of multimodal large language models to improve translation quality and evaluate the impact of translation unit size, context length, and propose 'token efficient' approach for manga translation.
Outcome: The proposed method achieves state-of-the-art results for Japanese-English translation and sets a new standard for Japanese and Polish translation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations