Challenge: Advertising banners are an instrumental medium in digital marketing campaigns.
Approach: They propose a training-free framework for fully automated banner ad design creation that enables frontier multimodal large language models to streamline the production of effective banners with minimal manual effort.
Outcome: The proposed framework is based on a training-free model that can be used to create fully automated banner ad design creations with minimal manual effort across diverse marketing contexts.

Similar Papers

Mirror in the Model: Ad Banner Image Generation via Reflective Multi-LLM and Multi-modal Agents (2025.emnlp-industry)

Copied to clipboard

Challenge: Recent advances in generative modeling have greatly improved image synthesis quality.
Approach: They propose an agentic refinement framework for automatic ad banner generation that integrates a hierarchical multimodal agent system with a coordination loop.
Outcome: The proposed model outperforms existing models in real-world banner design scenarios.
BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferences (2025.findings-emnlp)

Copied to clipboard

Challenge: Web banner advertisements are often selected manually because of human preferences . a new benchmark evaluates the degree of alignment with human preferences in two tasks .
Approach: a benchmark was developed to evaluate the human preference-driven banner selection process using vision-language models.
Outcome: The proposed benchmark assesses the degree of alignment with human preferences in two tasks using vision-language models.
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations (2025.acl-long)

Copied to clipboard

Challenge: State-of-the-art multimodal web agents can perform many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs).
Approach: They propose to build multimodal web agents for few-shot adaptability using human demonstrations to improve their generalization and adaptability.
Outcome: The proposed framework enables both proprietary and open-weights multimodal web agents to adapt to new websites and domains using few human demonstrations.
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity (2026.findings-eacl)

Copied to clipboard

Challenge: SMART-EDITOR is a framework for compositional layout and content editing for structured visual domains.
Approach: They introduce a framework for compositional editing for structured images like posters or websites . SMART-EDITOR maintains global coherence through two complementary strategies .
Outcome: The proposed framework maintains global coherence through two complementary strategies.
KAFA: Rethinking Image Ad Understanding with Knowledge-Augmented Feature Adaptation of Vision-Language Models (2023.acl-industry)

Copied to clipboard

Challenge: Image ad understanding is a crucial task with wide real-world applications, but is under-explored in the machine learning community due to the lack of foundational vision-language models (VLMs) .
Approach: They propose a simple feature adaptation strategy to fuse multimodal information for image ads and further empower it with knowledge of real-world entities.
Outcome: The proposed strategy fuses multimodal information for image ads and empowers it with knowledge of real-world entities.
The Revolution of Multimodal Large Language Models: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have led to the development of multimodal large language model.
Approach: They present a review of recent visual-based Large Language Models and analyze their architectures and alignment strategies.
Outcome: The proposed models can integrate visual and textual modalities while providing a dialogue-based interface and instruction-following capabilities.
ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs’ Capability via Chart Editing (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of multimodal large language models rely on limited case studies . however, they lack the ability to generate accurate edits according to the instructions .
Approach: They propose a benchmark for chart editing that includes 1,405 edit instructions applied to 233 real-world charts.
Outcome: The proposed benchmark includes 1,405 diverse editing instructions applied to 233 real-world charts.
Towards Unified Multimodal Large Language Models: A survey (2026.findings-acl)

Copied to clipboard

Challenge: unified multimodal large language models (MLLMs) are emerging but lack a systematic framework to connect them and situate current trends within a broader landscape.
Approach: They present a systematic review of unified Multimodal Large Language Models . they outline the foundational concepts and prerequisites for understanding them .
Outcome: The present review provides a systematic and systematic overview of unified MLLMs . it discusses persistent challenges and identify promising directions for future research .
Unveiling Multimodal Processing: Exploring Activation Patterns in Multimodal LLMs for Interpretability and Efficiency (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in multimodal large language models have remained opaque.
Approach: They propose a method to convert dense MLLMs into fine-grained Mixture-of-Experts architectures.
Outcome: The proposed method outperforms random expert pruning and sparse activation and model pruning.
Learning from LLM Agents: In-Context Generative Models for Text Casing in E-Commerce Ads (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing NER-based transformer models are expensive and lack contextual dependencies, making them less reliable when handling unseen or ad-specific terms, e.g., brand names.
Approach: They propose a two-stage approach to casing correction in e-commerce ad content that leverages Chain-of-Actions to enforce content policies while accurately handling ads-specific terms.
Outcome: The proposed model outperforms existing NER-based models and achieves near-LLM performance at a fraction of the cost.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations