Papers by Hongda Shen

2 papers
Visual Zero-Shot E-Commerce Product Attribute Value Extraction (2025.naacl-industry)

Copied to clipboard

Challenge: Existing zero-shot product attribute value extraction approaches require sellers to manually provide product descriptions.
Approach: They propose a cross-modal zero-shot attribute value generation framework based on CLIP that uses product images as inputs for zero- shot inference.
Outcome: The proposed framework significantly outperforms other vision-language models for zero-shot attribute value extraction.
MICE: Mixture of Image Captioning Experts Augmented e-Commerce Product Attribute Value Extraction (2025.acl-industry)

Copied to clipboard

Challenge: Existing visual attribute value extraction methods rely on product images and textual information, which can be ambiguous, inaccurate, or unavailable.
Approach: They propose a framework that leverages a curated pool of image captioning models to generate accurate captions from product images.
Outcome: The proposed framework significantly improves state-of-the-art large multimodal models in zero-shot and fine-tuning settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations