Papers by Hongda Shen
Visual Zero-Shot E-Commerce Product Attribute Value Extraction (2025.naacl-industry)
Copied to clipboard
| Challenge: | Existing zero-shot product attribute value extraction approaches require sellers to manually provide product descriptions. |
| Approach: | They propose a cross-modal zero-shot attribute value generation framework based on CLIP that uses product images as inputs for zero- shot inference. |
| Outcome: | The proposed framework significantly outperforms other vision-language models for zero-shot attribute value extraction. |
MICE: Mixture of Image Captioning Experts Augmented e-Commerce Product Attribute Value Extraction (2025.acl-industry)
Copied to clipboard
| Challenge: | Existing visual attribute value extraction methods rely on product images and textual information, which can be ambiguous, inaccurate, or unavailable. |
| Approach: | They propose a framework that leverages a curated pool of image captioning models to generate accurate captions from product images. |
| Outcome: | The proposed framework significantly improves state-of-the-art large multimodal models in zero-shot and fine-tuning settings. |