Detecting AI-Generated Video: A Vision–Language Dual-View Survey (2026.findings-acl)
Copied to clipboard
| Challenge: | realism of AI-generated Videos (AIGC-V) rendering artifact-centric detection insufficient, authors argue . a vision–language dual-view taxonomy is proposed to systematize this rapidly evolving field . |
| Approach: | They propose a Vision–Language Dual-View taxonomy to systematize AIGC-V detection . they propose realism of AI-generated Videos is rendering traditional inspection insufficient . |
| Outcome: | The proposed model aims to show that the existing methods are consistent with real-world facts. |
Similar Papers
Reasoning-Aware AIGC Detection via Alignment and Reinforcement (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to AIGC detection have relied on statistical classifiers or black-box neural models, which exploit surface-level patterns and struggle to generalize as LLMs evolve. |
| Approach: | They propose a framework that generates interpretable reasoning chains before classification using supervised fine-tuning and reinforcement learning to improve accuracy. |
| Outcome: | The proposed framework achieves state-of-the-art performance across multiple benchmarks, offering a robust and transparent solution for AIGC detection. |
Detecting AI-Generated Content on Social Media with Multi-modal Language Models (2026.acl-industry)
Copied to clipboard
Chenyang Yang, Shen Yan, Yibo Yang, Litao Hu, Yuchen Liu, Yuan Zeng, Hanchao Yu, Yinan Zhu, Sumedha Singla, Brian Vanover, Huijun Qian, Zihao Wang, Fujun Liu, Aashu Singh, Jianyu Wang, Xuewen Zhang
| Challenge: | Existing methods for AI-generated content detection face poor generalization to newer models, reliance on single modalities, and lack of interpretable explanations. |
| Approach: | They propose a model that curates diverse social media data and trains a vision-language model for detection and explanation. |
| Outcome: | The proposed model achieves state-of-the-art detection performance on public benchmarks and observes positive downstream impacts on user engagement. |
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? (2025.emnlp-main)
Copied to clipboard
Yiyang Zhou, Linjie Li, Shi Qiu, Zhengyuan Yang, Yuyang Zhao, Siwei Han, Yangfan He, Kangqi Li, Haonian Ji, Zihao Zhao, Haibo Tong, Lijuan Wang, Huaxiu Yao
| Challenge: | Existing video benchmarks often resemble image-based questions with scans of only a few key frames, without deep temporal reasoning. |
| Approach: | They propose a video benchmark to assess whether large vision-language models can genuinely think with videos rather than perform superficial frame-level analysis. |
| Outcome: | The proposed benchmark consists of 3,269 videos and over 4,342 highly visual-centric questions across 11 categories, including Trajectory Analysis, Temporal Reasoning, and Forensics Detection. |
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos (2025.acl-long)
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) are used for video quality assessment, image captioning and video analysis. |
| Approach: | They propose a benchmark to evaluate MLLMs on AIGC videos using coherence validation, error awareness, error type detection and reasoning evaluation tasks. |
| Outcome: | The proposed benchmark evaluates 13 frontier MLLMs on AIGC videos. |
A Survey on Detection of LLMs-Generated Content (2024.findings-emnlp)
Copied to clipboard
Xianjun Yang, Liangming Pan, Xuandong Zhao, Haifeng Chen, Linda Petzold, William Yang Wang, Wei Cheng
| Challenge: | Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance . |
| Approach: | They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy. |
| Outcome: | The proposed model will be able to detect human-written content in real time. |
Let’s Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought (2023.emnlp-main)
Copied to clipboard
Vaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He, Alex Mei, Yujie Lu, Chinmay Sonar, Michael Saxon, William Wang
| Challenge: | Existing studies show vision-language systems can reason about images using natural language, but their capacity for video reasoning remains underexplored. |
| Approach: | They propose to frame video reasoning as the sequential understanding of a small number of keyframes, thereby leveraging the power and robustness of vision-language systems' capacity to reason about images using natural language. |
| Outcome: | The proposed models can generate multiple intermediate keyframes and predict future keyframe, and they perform poorly on GPT-4, GPT-3, and VICUNA. |
From Detection to Understanding: Multi-Turn Reasoning for Video Misinformation Analysis (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on binary veracity judgments and do not evaluate process-level justifications for misinformation models. |
| Approach: | They propose a video misinformation analysis benchmark that assesses reasoning in video misinterpretation. |
| Outcome: | The proposed framework improves reasoning accuracy and explanation quality compared to existing models . it covers 12 fine-grained deception categories and progresses from perceptual attribution to intent and persuasion analysis. |
ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on question generation from videos are mostly focused on generating questions about common objects and attributes. |
| Approach: | They propose a model architecture combining Transformers, rich context signals and a combination of cross-entropy and contrastive loss function to encourage entity-centric question generation. |
| Outcome: | The proposed system yields BLEU, ROUGE, CIDEr, and METEOR scores of 71.3, 78.6, 7.31, and 81.9. |
Bringing Real-World Relations into Video Generation with Graph-Structured Knowledge (2026.acl-long)
Copied to clipboard
| Challenge: | Existing text-to-video models struggle to accurately simulate real-world physics and dynamic entity interactions. |
| Approach: | They propose a framework that integrates graph-structured temporal knowledge into video latent diffusion models to enhance compositional generation and interaction fidelity. |
| Outcome: | The proposed framework enhances compositional generation and interaction fidelity by integrating graph-structured temporal knowledge into video latent diffusion models. |
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation methods for complex multi-shot video are anchored to single-shot paradigms, lacking comprehensive story assets and cross-shot metrics. |
| Approach: | They propose a framework that synergizes the high-level semantic reasoning of Large Multimodal Models with the fine-grained perceptual rigor of domain-specific expert models. |
| Outcome: | The proposed framework synergizes the high-level semantic reasoning of Large Multimodal Models with the fine-grained perceptual rigor of domain-specific expert models. |