IPS: In-Prompt Process Supervision for Short Video Content Moderation (2026.acl-industry)
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) capture semantics of short video content but fail to account for policy-specific details. |
| Approach: | They propose a framework that integrates In-prompt Process Supervision into MLLMs . they propose sequential reasoning over ancillary questions during fine-tuning . |
| Outcome: | IPS outperforms baseline MLLMs on public and proprietary benchmarks . replacing human-annotated ancillary labels with MLML-generated ones results in performance degradation. |
Similar Papers
Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation (2025.acl-industry)
Copied to clipboard
Zixuan Wang, Jinghao Shi, Hanzhong Liang, Xiang Shen, Vera Wen, Zhiqian Chen, Yifan Wu, Zhixin Zhang, Hongyu Xiong
| Challenge: | Effective content moderation is essential for video platforms to safeguard user experience and uphold community standards. |
| Approach: | They propose a method to transform a generative MLLM into a multimodal classifier using minimal discriminative training data. |
| Outcome: | The proposed method improves F1 score by 66.50% over traditional classifiers while requiring only 2% of the fine-tuning data. |
Reasoning-Enhanced Domain-Adaptive Pretraining of Multimodal Large Language Models for Short Video Content Governance (2025.emnlp-industry)
Copied to clipboard
Zixuan Wang, Yu Sun, Hongwei Wang, Baoyu Jing, Xiang Shen, Xin Dong, Zhuolin Hao, Hongyu Xiong, Yang Song
| Challenge: | Existing approaches to identifying inappropriate content require extensive human-labeled data and lack cross-issue generalization. |
| Approach: | They propose a reasoning-enhanced multimodal large language model (MLLM) pretraining paradigm for unified inappropriate content detection. |
| Outcome: | The proposed model improves the MLLM's performance in both zero-shot and supervised fine-tuning settings and shows strong generalization capabilities to emergent, previously unseen issues. |
Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation (2026.findings-acl)
Copied to clipboard
| Challenge: | Short-video platforms have become major channels for misinformation, but their robustness against misinformation entangled with cognitive biases remains under-explored. |
| Approach: | They propose a framework for evaluation of short-video platforms that use visual cues and social cue. |
| Outcome: | The proposed framework evaluates MLLMs across five modality settings. |
WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing MCoT methods focus on inter-object reasoning, overlooking intra-object understanding crucial for image classification. |
| Approach: | They propose a Weak-supervision-guided Step-by-step Explanation method that reformulates MCoTs under weak supervision into concise, interpretable reasoning chains. |
| Outcome: | The proposed method improves interpretability by 37% and improves classification accuracy. |
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)
Copied to clipboard
Pei Fu, Tongkun Guan, Zining Wang, Zhentao Guo, Chen Duan, Hao Sun, Boming Chen, Qianyi Jiang, Jiayao Ma, Kai Zhou, Junfeng Luo
| Challenge: | Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms. |
| Approach: | They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks. |
| Outcome: | The proposed models perform well on mainstream benchmarks and are compared with other models. |
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos (2025.acl-long)
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) are used for video quality assessment, image captioning and video analysis. |
| Approach: | They propose a benchmark to evaluate MLLMs on AIGC videos using coherence validation, error awareness, error type detection and reasoning evaluation tasks. |
| Outcome: | The proposed benchmark evaluates 13 frontier MLLMs on AIGC videos. |
Towards Unified Multimodal Large Language Models: A survey (2026.findings-acl)
Copied to clipboard
| Challenge: | unified multimodal large language models (MLLMs) are emerging but lack a systematic framework to connect them and situate current trends within a broader landscape. |
| Approach: | They present a systematic review of unified Multimodal Large Language Models . they outline the foundational concepts and prerequisites for understanding them . |
| Outcome: | The present review provides a systematic and systematic overview of unified MLLMs . it discusses persistent challenges and identify promising directions for future research . |
Can’t See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs (2025.acl-long)
Copied to clipboard
Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang, Youliang Yuan, Pinjia He, Shuai Wang, Zhaopeng Tu
| Challenge: | Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. |
| Approach: | They propose a multimodal safety awareness benchmark to evaluate MLLMs across 29 safety scenarios with 1,500 carefully curated image-prompt pairs. |
| Outcome: | The proposed model is able to identify unsafe content and avoid over-sensitivity that can hinder helpfulness. |
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback (2024.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have influenced the development of video large multimodal models (VLMMs). |
| Approach: | They propose a method that integrates video descriptions as context into a multimodal AI system to enrich the understanding of video content. |
| Outcome: | Empirical evaluations show that the proposed approach outperforms existing approaches for video large multimodal models (VLMMs) |
Probing Multimodal Large Language Models for Global and Local Semantic Representations (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have focused on the ability of MLLMs to generate single tokens one by one, while lacking studies about how their representation vectors can encode global multimodal information. |
| Approach: | They propose to use image-caption corpus to train Multimodal Large Language Models (MLLMs) . they find that the topmost layers encode more global semantic information . |
| Outcome: | The proposed models can encode more global semantic information, rather than the topmost layers, and perform better on visual-language entailment tasks. |