Papers by Mingwei Shen
Misspelling Detection from Noisy Product Images (2020.coling-industry)
Copied to clipboard
| Challenge: | Existing spelling research has focused on advancement in misspelling correction . a single inadvertent or intentional misspeller can propagate to large amounts of inventory . |
| Approach: | They propose a method to automatically detect misspellings from product images . they curate a large corpus and define a rich set of features to validate the method . |
| Outcome: | The proposed method improves on a large corpus and an out-of-domain public dataset and improves by 20% in the F1 score. |
ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration (2025.emnlp-main)
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have shown impressive capabilities in vision-language understanding but their visual input remains fixed throughout the reasoning process. |
| Approach: | They propose a model-agnostic tree search algorithm tailored for vision-level reasoning that allows MLLMs to explore textual tokens while visual input remains fixed throughout reasoning process. |
| Outcome: | The proposed algorithm outperforms strong large models such as GPT-4o on high-resolution benchmarks and improves performance on a series of elaborate high-level benchmarks. |
Semantic matching for text classification with complex class descriptions (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for text classification support zero-shot learning but not both . Existing approaches do not support zero or few-shot, and are insufficient for complex classes . |
| Approach: | They propose a method which rapidly adapts from seen classes to new/unseen ones . they use labels and complex class descriptions to perform zero- and few-shot learning . |
| Outcome: | The proposed method beats baselines on complex class descriptions by 22.48% . it also improves zero-shot learning by 4.29% . |
An Explainable Toolbox for Evaluating Pre-trained Vision-Language Models (2022.emnlp-demos)
Copied to clipboard
| Challenge: | Existing studies evaluate VLP models by comparing the fine-tuned downstream task performance with the average downstream task accuracy. |
| Approach: | They propose a toolbox for evaluating Vision-Language Pretraining (VLP) models. |
| Outcome: | The proposed toolbox provides the preliminary datasets that deepen the image-texting ability of a VLP model. |