Challenge: Existing methods for question decomposition focus on unimodal language models, but question decomposing capability of Multimodal Large Language Models (MLLMs) has yet to be explored.
Approach: They propose a finetuning dataset and a training objective for selective decomposition to enhance the model's question decomposing capability.
Outcome: The proposed dataset shows that existing models struggle to produce high-quality sub-questions.

Similar Papers

Is a Question Decomposition Unit All We Need? (2022.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LMs) have achieved state-of-the-art performance on many NLP benchmarks.
Approach: They propose to decompose a hard question into simpler questions that are easier for models to answer.
Outcome: The proposed approach significantly improves model performance (24% for GPT3 and 29% for RoBERTa-SQuAD along with a symbolic calculator) by decomposing a hard question into simpler questions that are easier for models to answer.
Towards Unified Multimodal Large Language Models: A survey (2026.findings-acl)

Copied to clipboard

Challenge: unified multimodal large language models (MLLMs) are emerging but lack a systematic framework to connect them and situate current trends within a broader landscape.
Approach: They present a systematic review of unified Multimodal Large Language Models . they outline the foundational concepts and prerequisites for understanding them .
Outcome: The present review provides a systematic and systematic overview of unified MLLMs . it discusses persistent challenges and identify promising directions for future research .
Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: MLLMs have achieved significant breakthroughs in understanding across text and vision, but current models still face inconsistencies in reasoning outcomes.
Approach: They propose to evaluate multimodal large language models using a multimodal knowledge reasoning dataset to examine the extent of consistency degradation.
Outcome: The proposed evaluation tasks show that MLLMs are inefficient at integrating knowledge across modalities .
Generating Complex Question Decompositions in the Face of Distribution Shifts (2025.naacl-long)

Copied to clipboard

Challenge: Question decomposition has been found to improve large language models’ (LLMs) performance on complex question answering (QA) however, performance on the task remains dominated by supervised approaches, suggesting room for making LLMs better decomposers.
Approach: They propose to generate synthetic decomposition data with only five annotated examples by extending recent advances in using LLM-as-judge and for reranking in novel ways.
Outcome: The proposed approach generates synthetic decomposition data with only five examples over two benchmark datasets.
VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions (2025.acl-long)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) can grasp the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent.
Approach: They propose a framework for multimodal large language models to grasp the intention of a question and decompose it into a series of visual recognition sub-tasks to find out the answer.
Outcome: The proposed framework improves the accuracy of complex video-related questions by 29.6% and 17.2% on CVQA and the existing VQA datasets.
Probing Logical Reasoning of MLLMs in Scientific Diagrams (2025.emnlp-main)

Copied to clipboard

Challenge: logical reasoning is key to real-world applications like science education, environmental monitoring, and medical diagnostics.
Approach: They construct visual questions that follow seven structured templates with progressively more complex reasoning involved.
Outcome: The proposed models perform logical inferences based on visual information.
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluation methodologies for multimodal large language models are limited in evaluating objective queries without considering real-world user experiences.
Approach: They propose to evaluate multimodal large language models with per-sample criteria using potent MLLM as the judge.
Outcome: The proposed evaluation paradigm shows that it can be used to evaluate multimodal large language models with per-sample criteria.
Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks? (2025.findings-acl)

Copied to clipboard

Challenge: Existing models are unable to resolve references to abstract visual stimuli, such as color patches and color grids, but their pragmatic capabilities are still a challenge for state-of-the-art MLLMs.
Approach: They investigate whether multimodal large language models are able to resolve references to abstract visual stimuli, such as color patches and color grids, in a well-known reference resolution paradigm.
Outcome: The proposed model can resolve references to abstract visual stimuli in dyadic reference games.
Protecting multimodal large language models against misleading visualizations (2026.acl-long)

Copied to clipboard

Challenge: MLLMs are robust to misleading visualizations, i.e., charts that distort the underlying data, leading readers to draw inaccurate conclusions.
Approach: They propose to use table-based QA and redrawing the visualization to improve QA performance on misleading visualizations.
Outcome: The proposed methods improve MLLM question-answering accuracy on misleading visualizations without compromising accuracy on non-misleading ones.
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms.
Approach: They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks.
Outcome: The proposed models perform well on mainstream benchmarks and are compared with other models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations