Challenge: Multimodal manga analysis focuses on enhancing manga understanding with visual and textual features.
Approach: They propose a task to enhance manga understanding with visual and textual features by providing a shared semantic space for vision and language understanding.
Outcome: The proposed task provides a shared semantic space for vision and language understanding.

Similar Papers

Context-Informed Machine Translation of Manga using Multimodal Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Automated manga translation is a promising potential solution, but it is underdeveloped due to the need to incorporate visual elements into the translation process to resolve ambiguities.
Approach: They propose a method that leverages the vision component of multimodal large language models to improve translation quality and evaluate the impact of translation unit size, context length, and propose 'token efficient' approach for manga translation.
Outcome: The proposed method achieves state-of-the-art results for Japanese-English translation and sets a new standard for Japanese and Polish translation.
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding (2026.findings-eacl)

Copied to clipboard

Challenge: Manga is a richly multimodal narrative form that blends images and text in complex ways.
Approach: They propose two benchmarks for multimodal manga understanding: mangaOCR and mangaVQA . mangaVQ consists of 526 high-quality, manually constructed question-answer pairs .
Outcome: The proposed model is finetuned from the open-source LMM Qwen2.5-VL . it compares with proprietary models such as GPT-4o and Gemini 2.5 to evaluate its performance .
ComicVQA: A Benchmark for Visual Reasoning in Multimodal LLMs (2026.findings-acl)

Copied to clipboard

Challenge: ComicVQA is a visual reasoning benchmark for comics.
Approach: They propose a comics-based benchmark for evaluating MLLMs on visual reasoning.
Outcome: The proposed model achieves 62.6% accuracy on Missing Panel Prediction and 46.4% on Panel Sorting, compared to open-source models.
Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text (2026.findings-eacl)

Copied to clipboard

Challenge: Existing multimodal models rely on machine translation, but performance drops for other languages due to limited multilingual multimodal resources.
Approach: They propose a lightweight alignment method that learns only a few linear layers using English text alone to map multilingual text embeddings into multimodal space.
Outcome: M2M achieves strong zero-shot transfer on XTD Text-to-Image retrieval in English and spanish . it learns only a few linear layers to map multilingual text embeddings into multimodal space .
MM-GATBT: Enriching Multimodal Representation Using Graph Attention Network (2022.naacl-srw)

Copied to clipboard

Challenge: Existing models that use a self-attention mechanism to create graphs with multiple modes ignore interaction between entities, multimodalities, or both.
Approach: They propose a multimodal graph representation learning model that captures relational semantics within one modality and interactions between different modalities.
Outcome: The proposed model outperforms existing models on the MM-IMDb dataset in all aspects of multimodal representation.
Exploring the Capability of Multimodal LLMs with Yonkoma Manga: The YManga Dataset and Its Challenging Tasks (2024.findings-emnlp)

Copied to clipboard

Challenge: YManga dataset is the first specifically designed for yonkoma manga understanding .
Approach: They propose to use a dataset of 1,015 yonkoma strips with 10,150 human annotations to define three tasks for panel sequence detection, intent generation and description generation for masked panels.
Outcome: The proposed dataset contains 1,015 high-quality yonkoma strips with 10,150 human annotations.
MM-LLMs: Recent Advances in MultiModal Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: MultiModal Large Language Models (MM-LLMs) have undergone significant advances in the past year . traditional MM models incur substantial computational costs, especially when trained from scratch .
Approach: They propose a taxonomy encompassing 126 MM-LLMs and summarize key training recipes to enhance their potency.
Outcome: The proposed models preserve the reasoning and decision-making capabilities of LLMs and empower diverse range of MM tasks.
LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine Translation (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances struggle to train a separate model for each language pair, which is costly and unaffordable when the number of languages increases in the real world.
Approach: They propose to train different MMT models to support translations between different languages.
Outcome: The proposed model is able to handle the above issues by providing a shared semantic space for multiple languages.
Domain Adaptation of Image Encoder for Multimodal Manga Translation (2026.eacl-srw)

Copied to clipboard

Challenge: Existing machine translation systems lack sufficient manga comprehension capabilities when utilizing image information.
Approach: They propose a domain-adapted image encoder training method for manga . the method trains encoders to acquire visual features that consider the structural and sequential characteristics of the manga based on a Japanese-English translation task.
Outcome: The proposed method improves translation evaluation metrics in Japanese-English translation task compared to the conventional method .
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems (2024.findings-emnlp)

Copied to clipboard

Challenge: Programming often involves translating detailed and complex specifications into code . current state-of-the-art models struggle to solve these problems, a new study shows .
Approach: They propose a multi-modal coding dataset to evaluate algorithmic problem-solving skills in visually rich contexts.
Outcome: The proposed model lacks powerful vision-code models due to the extreme demand for reasoning abilities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations