Vishnu Prabhakaran, Purav Aggarwal, Vishruit Kulshreshtha, Arunita Das, Sahini Venkata Sitaram Sruti, Anoop Saladi
| Challenge: | general-purpose vision-language models struggle to understand and converse about real-world e-commerce product images. |
| Approach: | a new approach is proposed to use large-scale image-text pairs to train a generative VLM for e-commerce product images. |
| Outcome: | The proposed model outperforms general-purpose VLMs on multiple vision tasks in the e-commerce domain. |
Similar Papers
ViFT: Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Visual instruction tuning is the predominant technology in eliciting multimodal task-solving capabilities of large vision-language models. |
| Approach: | They propose a visual instruction-free fine-tuning framework for large vision-language models . they require only text-only instructions and image caption data during training . |
| Outcome: | The proposed framework is based on visual instruction tuning, but requires images as input . it can achieve state-of-the-art performance on several downstream benchmarks with less training data. |
Adapting Vision-Language Models for E-commerce Understanding at Scale (2026.eacl-industry)
Copied to clipboard
Matteo Nulli, Vladimir Orshulevich, Tala Bazazo, Christian Herold, Michael Kozielski, Marcin Mazur, Szymon Tuzel, Cees G. M. Snoek, Seyyed Hadi Hashemi, Omar Javed, Yannick Versley, Shahram Khadivi
| Challenge: | Existing approaches to adapt VLMs to attribute-centric, multi-image, and noisy data are limited. |
| Approach: | They propose a novel evaluation suite that incorporates deep product understanding, strict instruction following, and dynamic attribute extraction. |
| Outcome: | The proposed model improves e-commerce performance while preserving broad multimodal capabilities. |
Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data (2024.findings-acl)
Copied to clipboard
Yanda Li, Chi Zhang, Gang Yu, Wanqi Yang, Zhibin Wang, Bin Fu, Guosheng Lin, Chunhua Shen, Ling Chen, Yunchao Wei
| Challenge: | OpenAI's GPT-4 has demonstrated remarkable multimodal capabilities, but specific mechanics of GPT4 remain unknown. |
| Approach: | They propose a data collection methodology that synchronously synthesizes images and dialogues for visual instruction tuning. |
| Outcome: | The proposed method improves on ten commonly assessed models and provides greater flexibility compared to existing methods. |
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning (2024.findings-acl)
Copied to clipboard
Zhiyang Xu, Chao Feng, Rulin Shao, Trevor Ashby, Ying Shen, Di Jin, Yu Cheng, Qifan Wang, Lifu Huang
| Challenge: | Recent vision-language models (VLMs) have shown impressive capabilities as general visual assistants, but there are two challenges to their performance: (1) lacking task diversity in pretraining and visual instruction tuning; (2) annotation error and bias in GPT-4 synthesized instruction tuning data. |
| Approach: | They propose a two-stage instruction tuning framework that fine tunes VLMs firstly and further tuned on GPT-4 synthesized data. |
| Outcome: | The proposed framework outperforms the traditional single-stage visual instruction tuning framework and achieves state-of-the-art performance across a wide range of multi-modal evaluation benchmarks. |
Visually-augmented pretrained language models for NLP tasks without images (2023.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to improve pre-trained language models lack visual commonsense and semantics. |
| Approach: | They propose a visual-augmented approach to fine-tune pre-trained language models by using retrieved or generated images instead of relying on explicit images. |
| Outcome: | The proposed approach outperforms baselines on ten tasks and consistently outperformed other approaches. |
Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection (2024.findings-acl)
Copied to clipboard
Ruibo Chen, Yihan Wu, Lichang Chen, Guodong Liu, Qi He, Tianyi Xiong, Chenxi Liu, Junfeng Guo, Heng Huang
| Challenge: | Existing data selection methods for instruction-following large language models rely on unreliable scores or use downstream tasks for selection. |
| Approach: | They propose a method that utilizes the VLM itself as a filter to select high-quality instruction-tuning data. |
| Outcome: | The proposed method can reach better results compared to full data settings with merely about 15% samples and can achieve superior performance against competitive baselines. |
VLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a significant drawback of Vision-language Models is their reliance on static training data, leading to outdated information and limited contextual awareness. |
| Approach: | They propose a framework with knowledge-enhanced reranking and noise-injected training to improve the VLM's ranking ability. |
| Outcome: | The proposed framework is based on a simple yet effective instruction template and is able to induce its ranking ability and serve it as a reranker to precisely filter the top-k retrieved images. |
Efficient Vision-Language pre-training via domain-specific learning for human activities (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current vision-language models owe their success to large-scale pretraining on web-collected data. |
| Approach: | They propose a domain-aligned pretraining strategy that aligns the downstream tasks to the downstream domain without additional data collection. |
| Outcome: | The proposed method outperforms existing models on large-scale vision-language training datasets while preserving generalist knowledge. |
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for generating product descriptions from images are inaccurate and generic . e-commerce product descriptions are important for content marketing and increasing engagement . |
| Approach: | They propose a new setting for generating product descriptions from images, augmented by marketing keywords. |
| Outcome: | The proposed approach improves the accuracy and diversity of product descriptions by up to 3.3% on Rouge-L and 9.4% on D-5. |
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation (2025.acl-long)
Copied to clipboard
Yue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta, Luca Weihs, Andrew Head, Mark Yatskar, Chris Callison-Burch, Ranjay Krishna, Aniruddha Kembhavi, Christopher Clark
| Challenge: | Vision-language models struggle to understand text-rich images due to the scarcity of diverse text-only large language data. |
| Approach: | They propose a framework that leverages the coding capabilities of text-only large language models to create synthetic text-rich multimodal data. |
| Outcome: | The proposed framework can generate high-quality instruction-tuning data using Python, HTML, LaTeX and other languages. |