Challenge: logical reasoning is key to real-world applications like science education, environmental monitoring, and medical diagnostics.
Approach: They construct visual questions that follow seven structured templates with progressively more complex reasoning involved.
Outcome: The proposed models perform logical inferences based on visual information.

Similar Papers

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Current scientific reasoning models struggle with generalization across domains and fall short of multimodal perception.
Approach: They propose to use multimodal large language models to integrate text, images, and other modalities to enhance scientific reasoning.
Outcome: The proposed models can integrate text, images, and other modalities and improve reasoning across disciplines.
The Revolution of Multimodal Large Language Models: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have led to the development of multimodal large language model.
Approach: They present a review of recent visual-based Large Language Models and analyze their architectures and alignment strategies.
Outcome: The proposed models can integrate visual and textual modalities while providing a dialogue-based interface and instruction-following capabilities.
Multimodal Causal Reasoning Benchmark: Challenging Multimodal Large Language Models to Discern Causal Links Across Modalities (2025.findings-acl)

Copied to clipboard

Challenge: Existing MLLMs lack robustness in multimodal causal reasoning compared to their performance in textual settings.
Approach: They propose a novel multimodal chain-of-thought (CoT) reasoning benchmark that leverages siamese images and text pairs to challenge MLLMs.
Outcome: The proposed benchmark leverages siamese images and text pairs to challenge MLLMs.
Probing Multimodal Large Language Models for Global and Local Semantic Representations (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have focused on the ability of MLLMs to generate single tokens one by one, while lacking studies about how their representation vectors can encode global multimodal information.
Approach: They propose to use image-caption corpus to train Multimodal Large Language Models (MLLMs) . they find that the topmost layers encode more global semantic information .
Outcome: The proposed models can encode more global semantic information, rather than the topmost layers, and perform better on visual-language entailment tasks.
Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: MLLMs have achieved significant breakthroughs in understanding across text and vision, but current models still face inconsistencies in reasoning outcomes.
Approach: They propose to evaluate multimodal large language models using a multimodal knowledge reasoning dataset to examine the extent of consistency degradation.
Outcome: The proposed evaluation tasks show that MLLMs are inefficient at integrating knowledge across modalities .
From Multimodal LLM to Human-level AI: Modality, Instruction, Reasoning, Efficiency and beyond (2024.lrec-tutorials)

Copied to clipboard

Challenge: This tutorial aims to deliver a comprehensive review of cutting-edge research in MLLMs.
Approach: This tutorial will review cutting-edge research in MLLMs and examine the impact of ML in learning and reasoning.
Outcome: This course will review cutting-edge research in MLLMs and examine the impact of ML models on learning, learning, and multimodal reasoning.
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms.
Approach: They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks.
Outcome: The proposed models perform well on mainstream benchmarks and are compared with other models.
Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks? (2025.findings-acl)

Copied to clipboard

Challenge: Existing models are unable to resolve references to abstract visual stimuli, such as color patches and color grids, but their pragmatic capabilities are still a challenge for state-of-the-art MLLMs.
Approach: They investigate whether multimodal large language models are able to resolve references to abstract visual stimuli, such as color patches and color grids, in a well-known reference resolution paradigm.
Outcome: The proposed model can resolve references to abstract visual stimuli in dyadic reference games.
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models struggle with visual reasoning, despite strong performance on vision-language tasks.
Approach: They propose a visually cued chain-of-thought prompting that enhances multi-step mathematical reasoning by explicitly referencing visual annotations in diagrams.
Outcome: The proposed model improves GPT-4o's accuracy on an irregular polygon side-counting task from 7% to 93%.
Towards Unified Multimodal Large Language Models: A survey (2026.findings-acl)

Copied to clipboard

Challenge: unified multimodal large language models (MLLMs) are emerging but lack a systematic framework to connect them and situate current trends within a broader landscape.
Approach: They present a systematic review of unified Multimodal Large Language Models . they outline the foundational concepts and prerequisites for understanding them .
Outcome: The present review provides a systematic and systematic overview of unified MLLMs . it discusses persistent challenges and identify promising directions for future research .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations