| Challenge: | Multimodal large language models are increasingly used for movie understanding . however, their performance on movies lags behind other video understanding tasks . |
| Approach: | They analyze movie knowledge, cinematographic knowledge, and critical analysis to identify where MLLMs fail . ML models are increasingly used for movie understanding . |
| Outcome: | The results show that MLLMs outperform existing methods in small-scale settings involving factual knowledge but fail when cinematographic and critical analysis is required. |
Similar Papers
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)
Copied to clipboard
Pei Fu, Tongkun Guan, Zining Wang, Zhentao Guo, Chen Duan, Hao Sun, Boming Chen, Qianyi Jiang, Jiayao Ma, Kai Zhou, Junfeng Luo
| Challenge: | Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms. |
| Approach: | They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks. |
| Outcome: | The proposed models perform well on mainstream benchmarks and are compared with other models. |
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have demonstrated remarkable performance across various tasks, effectively following instructions to meet diverse user needs. |
| Approach: | They propose a framework for evaluation benchmarks and attack techniques for LLMs and MLLMs to enhance their security. |
| Outcome: | The proposed frameworks have been exploited to exploit the weaknesses of LLMs and MLLMs. |
MM-LLMs: Recent Advances in MultiModal Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | MultiModal Large Language Models (MM-LLMs) have undergone significant advances in the past year . traditional MM models incur substantial computational costs, especially when trained from scratch . |
| Approach: | They propose a taxonomy encompassing 126 MM-LLMs and summarize key training recipes to enhance their potency. |
| Outcome: | The proposed models preserve the reasoning and decision-making capabilities of LLMs and empower diverse range of MM tasks. |
Probing Multimodal Large Language Models for Global and Local Semantic Representations (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have focused on the ability of MLLMs to generate single tokens one by one, while lacking studies about how their representation vectors can encode global multimodal information. |
| Approach: | They propose to use image-caption corpus to train Multimodal Large Language Models (MLLMs) . they find that the topmost layers encode more global semantic information . |
| Outcome: | The proposed models can encode more global semantic information, rather than the topmost layers, and perform better on visual-language entailment tasks. |
Towards Unified Multimodal Large Language Models: A survey (2026.findings-acl)
Copied to clipboard
| Challenge: | unified multimodal large language models (MLLMs) are emerging but lack a systematic framework to connect them and situate current trends within a broader landscape. |
| Approach: | They present a systematic review of unified Multimodal Large Language Models . they outline the foundational concepts and prerequisites for understanding them . |
| Outcome: | The present review provides a systematic and systematic overview of unified MLLMs . it discusses persistent challenges and identify promising directions for future research . |
Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Multimodal large language models combine visual and textual data for tasks like image captioning and visual question answering. |
| Approach: | They propose temperature scaling and iterative prompt optimization to calibrate MLLMs and enhance model reliability. |
| Outcome: | The proposed techniques improve MLLMs and improve model reliability. |
Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing models are unable to resolve references to abstract visual stimuli, such as color patches and color grids, but their pragmatic capabilities are still a challenge for state-of-the-art MLLMs. |
| Approach: | They investigate whether multimodal large language models are able to resolve references to abstract visual stimuli, such as color patches and color grids, in a well-known reference resolution paradigm. |
| Outcome: | The proposed model can resolve references to abstract visual stimuli in dyadic reference games. |
From Multimodal LLM to Human-level AI: Modality, Instruction, Reasoning, Efficiency and beyond (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | This tutorial aims to deliver a comprehensive review of cutting-edge research in MLLMs. |
| Approach: | This tutorial will review cutting-edge research in MLLMs and examine the impact of ML in learning and reasoning. |
| Outcome: | This course will review cutting-edge research in MLLMs and examine the impact of ML models on learning, learning, and multimodal reasoning. |
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria (2025.naacl-long)
Copied to clipboard
Wentao Ge, Shunian Chen, Hardy Chen, Nuo Chen, Junying Chen, Zhihong Chen, Wenya Xie, Shuo Yan, ChenghaoZhu ChenghaoZhu, Ziyue Lin, Dingjie Song, Xidong Wang, Anningzhe Gao, Zhang Zhiyi, Jianquan Li, Xiang Wan, Benyou Wang
| Challenge: | Existing evaluation methodologies for multimodal large language models are limited in evaluating objective queries without considering real-world user experiences. |
| Approach: | They propose to evaluate multimodal large language models with per-sample criteria using potent MLLM as the judge. |
| Outcome: | The proposed evaluation paradigm shows that it can be used to evaluate multimodal large language models with per-sample criteria. |
Self-Improvement in Multimodal Large Language Models: A Survey (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Using data and data, self-improvement for Large Language Models has improved model capabilities without significantly increasing costs. |
| Approach: | This survey provides a comprehensive overview of self-improvement for Large Language Models . it includes commonly used evaluations and downstream applications . |
| Outcome: | The authors provide a comprehensive overview of self-improvement in Multimodal LLMs. |