BannerAgency: Advertising Banner Design with Multimodal LLM Agents (2025.emnlp-main)
Copied to clipboard
| Challenge: | Advertising banners are an instrumental medium in digital marketing campaigns. |
| Approach: | They propose a training-free framework for fully automated banner ad design creation that enables frontier multimodal large language models to streamline the production of effective banners with minimal manual effort. |
| Outcome: | The proposed framework is based on a training-free model that can be used to create fully automated banner ad design creations with minimal manual effort across diverse marketing contexts. |
Similar Papers
Mirror in the Model: Ad Banner Image Generation via Reflective Multi-LLM and Multi-modal Agents (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Recent advances in generative modeling have greatly improved image synthesis quality. |
| Approach: | They propose an agentic refinement framework for automatic ad banner generation that integrates a hierarchical multimodal agent system with a coordination loop. |
| Outcome: | The proposed model outperforms existing models in real-world banner design scenarios. |
BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferences (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Web banner advertisements are often selected manually because of human preferences . a new benchmark evaluates the degree of alignment with human preferences in two tasks . |
| Approach: | a benchmark was developed to evaluate the human preference-driven banner selection process using vision-language models. |
| Outcome: | The proposed benchmark assesses the degree of alignment with human preferences in two tasks using vision-language models. |
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations (2025.acl-long)
Copied to clipboard
| Challenge: | State-of-the-art multimodal web agents can perform many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs). |
| Approach: | They propose to build multimodal web agents for few-shot adaptability using human demonstrations to improve their generalization and adaptability. |
| Outcome: | The proposed framework enables both proprietary and open-weights multimodal web agents to adapt to new websites and domains using few human demonstrations. |
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity (2026.findings-eacl)
Copied to clipboard
| Challenge: | SMART-EDITOR is a framework for compositional layout and content editing for structured visual domains. |
| Approach: | They introduce a framework for compositional editing for structured images like posters or websites . SMART-EDITOR maintains global coherence through two complementary strategies . |
| Outcome: | The proposed framework maintains global coherence through two complementary strategies. |
KAFA: Rethinking Image Ad Understanding with Knowledge-Augmented Feature Adaptation of Vision-Language Models (2023.acl-industry)
Copied to clipboard
| Challenge: | Image ad understanding is a crucial task with wide real-world applications, but is under-explored in the machine learning community due to the lack of foundational vision-language models (VLMs) . |
| Approach: | They propose a simple feature adaptation strategy to fuse multimodal information for image ads and further empower it with knowledge of real-world entities. |
| Outcome: | The proposed strategy fuses multimodal information for image ads and empowers it with knowledge of real-world entities. |
The Revolution of Multimodal Large Language Models: A Survey (2024.findings-acl)
Copied to clipboard
Davide Caffagni, Federico Cocchi, Luca Barsellotti, Nicholas Moratelli, Sara Sarto, Lorenzo Baraldi, Lorenzo Baraldi, Marcella Cornia, Rita Cucchiara
| Challenge: | Recent advances in large language models have led to the development of multimodal large language model. |
| Approach: | They present a review of recent visual-based Large Language Models and analyze their architectures and alignment strategies. |
| Outcome: | The proposed models can integrate visual and textual modalities while providing a dialogue-based interface and instruction-following capabilities. |
ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs’ Capability via Chart Editing (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations of multimodal large language models rely on limited case studies . however, they lack the ability to generate accurate edits according to the instructions . |
| Approach: | They propose a benchmark for chart editing that includes 1,405 edit instructions applied to 233 real-world charts. |
| Outcome: | The proposed benchmark includes 1,405 diverse editing instructions applied to 233 real-world charts. |
Towards Unified Multimodal Large Language Models: A survey (2026.findings-acl)
Copied to clipboard
| Challenge: | unified multimodal large language models (MLLMs) are emerging but lack a systematic framework to connect them and situate current trends within a broader landscape. |
| Approach: | They present a systematic review of unified Multimodal Large Language Models . they outline the foundational concepts and prerequisites for understanding them . |
| Outcome: | The present review provides a systematic and systematic overview of unified MLLMs . it discusses persistent challenges and identify promising directions for future research . |
Unveiling Multimodal Processing: Exploring Activation Patterns in Multimodal LLMs for Interpretability and Efficiency (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in multimodal large language models have remained opaque. |
| Approach: | They propose a method to convert dense MLLMs into fine-grained Mixture-of-Experts architectures. |
| Outcome: | The proposed method outperforms random expert pruning and sparse activation and model pruning. |
Learning from LLM Agents: In-Context Generative Models for Text Casing in E-Commerce Ads (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Existing NER-based transformer models are expensive and lack contextual dependencies, making them less reliable when handling unseen or ad-specific terms, e.g., brand names. |
| Approach: | They propose a two-stage approach to casing correction in e-commerce ad content that leverages Chain-of-Actions to enforce content policies while accurately handling ads-specific terms. |
| Outcome: | The proposed model outperforms existing NER-based models and achieves near-LLM performance at a fraction of the cost. |