Challenge: Existing VLMs perform well on general multimodal tasks, but limited labeled data makes them difficult to apply to real-world business decisions.
Approach: They propose a new task that aims to rank ads for a target brand prior to deployment . they propose 'brand-specific ad ranking' which uses brand-specific effectiveness .
Outcome: The proposed task outperforms baselines on 10 brands on real-world advertising data.

Similar Papers

Learning from LLM Agents: In-Context Generative Models for Text Casing in E-Commerce Ads (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing NER-based transformer models are expensive and lack contextual dependencies, making them less reliable when handling unseen or ad-specific terms, e.g., brand names.
Approach: They propose a two-stage approach to casing correction in e-commerce ad content that leverages Chain-of-Actions to enforce content policies while accurately handling ads-specific terms.
Outcome: The proposed model outperforms existing NER-based models and achieves near-LLM performance at a fraction of the cost.
BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferences (2025.findings-emnlp)

Copied to clipboard

Challenge: Web banner advertisements are often selected manually because of human preferences . a new benchmark evaluates the degree of alignment with human preferences in two tasks .
Approach: a benchmark was developed to evaluate the human preference-driven banner selection process using vision-language models.
Outcome: The proposed benchmark assesses the degree of alignment with human preferences in two tasks using vision-language models.
VLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training (2025.findings-emnlp)

Copied to clipboard

Challenge: a significant drawback of Vision-language Models is their reliance on static training data, leading to outdated information and limited contextual awareness.
Approach: They propose a framework with knowledge-enhanced reranking and noise-injected training to improve the VLM's ranking ability.
Outcome: The proposed framework is based on a simple yet effective instruction template and is able to induce its ranking ability and serve it as a reranker to precisely filter the top-k retrieved images.
Top-Rank-Focused Adaptive Vote Collection for the Evaluation of Domain-Specific Semantic Models (2020.emnlp-main)

Copied to clipboard

Challenge: Embedding-based models are increasingly needed for domain-specific evaluation datasets.
Approach: They propose a protocol for the construction of a relatedness-based evaluation dataset based on adaptive pairwise comparisons and appropriate metrics to evaluate a semantic model via the aforementioned dataset.
Outcome: The proposed protocol is particularly accurate in top-rank evaluation.
KAFA: Rethinking Image Ad Understanding with Knowledge-Augmented Feature Adaptation of Vision-Language Models (2023.acl-industry)

Copied to clipboard

Challenge: Image ad understanding is a crucial task with wide real-world applications, but is under-explored in the machine learning community due to the lack of foundational vision-language models (VLMs) .
Approach: They propose a simple feature adaptation strategy to fuse multimodal information for image ads and further empower it with knowledge of real-world entities.
Outcome: The proposed strategy fuses multimodal information for image ads and empowers it with knowledge of real-world entities.
From Mimesis to Metamorphosis: Evolving VLM Judges via In-Context Comparing and Knowledge Internalization (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to subjective assessment are inconsistent and inconsistent due to inconsistent scales and inherent preference biases.
Approach: They propose a framework that operationalizes subjective assessment as comparative analysis and internalizes it via Language Buttons.
Outcome: The proposed framework achieves state-of-the-art performance across multiple benchmarks and is scale-steerable.
CARES: Context-Aware Resolution Selector for VLMs (2026.acl-long)

Copied to clipboard

Challenge: Large vision–language models process images at native or high resolution to remain effective across tasks.
Approach: They propose a lightweight preprocessing module that predicts the minimum sufficient input resolution for large vision–language models.
Outcome: CARES predicts when a pre-trained VLM's response converges to its peak ability to answer correctly, reducing compute by up to 80%.
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Existing metrics for long-form text outputs are prone to biases and scaling up is expensive.
Approach: They propose to evaluate VLMs with VLM feedback dataset . they use 15K customized score rubrics to train Prometheus-Vision .
Outcome: The proposed model shows highest correlation with human evaluators and GPT-4V among open-source models.
Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity (2025.emnlp-main)

Copied to clipboard

Challenge: Evaluating creativity is challenging, even for humans, because of its subjectivity and complex cognitive processes.
Approach: They propose a set of tasks to break down visual advertisement creativity into atypicality and originality with fine-grained annotations by humans.
Outcome: The proposed tasks demonstrate the promise and challenges of using VLMs for automated creativity assessment.
AdDriftBench: A Benchmark for Detecting Data Drift and Label Drift in Short Video Advertising (2025.findings-emnlp)

Copied to clipboard

Challenge: Short video advertising scenarios present unique challenges due to data drift (DD) and label drift (LD).
Approach: They propose to use data drift and label drift to evaluate models under rapidly shifting content distributions and labeling scenarios to assess their generalization capabilities.
Outcome: The proposed model performs moderately in short video advertising contexts, particularly in handling fine-grained semantics and adapting to shifting instructions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations