AdDriftBench: A Benchmark for Detecting Data Drift and Label Drift in Short Video Advertising (2025.findings-emnlp)
Copied to clipboard
Yinghao Song, Xiangji Zeng, Shuai Cui, Lu Sun, Zhaowei Liu, Yuan Yuan, Yulu Wang, Hai Zhou, Zhaohan Gong
| Challenge: | Short video advertising scenarios present unique challenges due to data drift (DD) and label drift (LD). |
| Approach: | They propose to use data drift and label drift to evaluate models under rapidly shifting content distributions and labeling scenarios to assess their generalization capabilities. |
| Outcome: | The proposed model performs moderately in short video advertising contexts, particularly in handling fine-grained semantics and adapting to shifting instructions. |
Similar Papers
NLP-ADBench: NLP Anomaly Detection Benchmark (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Anomaly detection (AD) is an important machine learning task, but its effectiveness in detecting harmful content, phishing attempts, and spam reviews is limited. |
| Approach: | They introduce NLP-ADBench, the most comprehensive NLP anomaly detection benchmark to date . it includes eight curated datasets and 19 state-of-the-art algorithms . |
| Outcome: | The NLP-ADBench benchmark includes 19 state-of-the-art methods and 8 curated datasets . no single model dominates across all datasets, indicating need for automated model selection . |
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for accelerating Large Vision-Language Models lack comprehensive evaluation across diverse backbones, benchmarks, and metrics. |
| Approach: | They propose EffiVLM-BENCH framework for evaluating absolute performance and generalization and loyalty. |
| Outcome: | The proposed framework offers insights into optimal strategies for accelerating LVLMs. |
MVTamperBench: Evaluating Robustness of Vision-Language Models (2025.findings-acl)
Copied to clipboard
Amit Agarwal, Srikant Panda, Angeline Charles, Hitesh Laxmichand Patel, Bhargava Kumar, Priyaranjan Pattnayak, Taki Hasan Rafi, Tejaswini Kumar, Hansa Meghwani, Karan Gupta, Dong-Kyu Chae
| Challenge: | Multimodal Large Language Models (MLLMs) have been a key advance in video understanding but their vulnerability to adversarial tampering remains underexplored. |
| Approach: | They evaluate MLLMs against five prevalent tampering techniques to assess their robustness . they use a tampered video format to examine the vulnerability of ML models . |
| Outcome: | The benchmark evaluates MLLMs against five prevalent tampering techniques based on 19 video manipulation tasks. |
Striking Gold in Advertising: Standardization and Exploration of Ad Text Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks and problem sets for automatic ad text generation are lacking . however, the growing volume of search queries has fueled research on the automatic generation of ads. |
| Approach: | They propose to standardize the task of automatic ad text generation (ATG) using a benchmark dataset, CAMERA, to enable the utilization of multi-modal information and facilitate industry-wise evaluations. |
| Outcome: | The proposed dataset standardizes the task of automatic ad text generation (ATG) it shows that existing metrics align with human evaluations and that the proposed methods can be used to improve the quality of the results. |
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on basic abilities using nonverbal methods, such as yes-no and multiple-choice questions. |
| Approach: | They propose a benchmark that provides more nuanced evaluations of alignment capabilities for large Vision-Language Models (VLMs) they use a rule-calibrated evaluator that exceeds GPT-4's evaluation ability and a “alignment score” to assess the robustness and stability of models across diverse prompts. |
| Outcome: | The proposed benchmark covers 13 tasks across three categories and includes both single-turn and multi-turn dialogue scenarios. |
FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing RAG research focuses on textual data, overlooking rich visual content in financial documents. |
| Approach: | They propose a visual RAG benchmark tailored for finance that integrates multimodal data and provides visual citation to ensure traceability. |
| Outcome: | The proposed visual RAG benchmark integrates multimodal data and provides visual citation to ensure traceability. |
BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferences (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Web banner advertisements are often selected manually because of human preferences . a new benchmark evaluates the degree of alignment with human preferences in two tasks . |
| Approach: | a benchmark was developed to evaluate the human preference-driven banner selection process using vision-language models. |
| Outcome: | The proposed benchmark assesses the degree of alignment with human preferences in two tasks using vision-language models. |
Dynabench: Rethinking Benchmarking in NLP (2021.naacl-main)
Copied to clipboard
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams
| Challenge: | Dynabench is an open-source platform for dynamic dataset creation and model benchmarking. |
| Approach: | They propose an open-source platform for dynamic dataset creation and model benchmarking. |
| Outcome: | The proposed platform can be used to create models that fail on simple challenges and falter in real-world scenarios. |
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Current sycophancy research has largely overlooked its specific manifestations in the video-language domain. |
| Approach: | They propose a video-LLM sycophancy benchmarking and evaluation to evaluate scophancies in video-LLMs. |
| Outcome: | The proposed benchmark evaluates sycophantic behavior in state-of-the-art Video-LLMs across diverse question formats, prompt biases, and visual reasoning tasks. |
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation (2025.findings-acl)
Copied to clipboard
Xinlong Chen, Yuanxing Zhang, Chongling Rao, Yushuo Guan, Jiaheng Liu, Fuzheng Zhang, Chengru Song, Qiang Liu, Di Zhang, Tieniu Tan
| Challenge: | Existing studies have not identified a link between video caption evaluation and T2V generation. |
| Approach: | They propose a video caption evaluation scheme specifically designed for T2V generation that integrates video annotation with caption evaluation. |
| Outcome: | The proposed system is agnostic to any particular caption format and can be used for training. |