Papers by Ran Jiao
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models (2025.emnlp-main)
Copied to clipboard
Qihang Ma, Shengyu Li, Jie Tang, Dingkang Yang, null Chenshaodong, Yingyi Zhang, Chao Feng, Ran Jiao
| Challenge: | Multi-modal keyphrase prediction (MMKP) aims to produce concise, informative phrases that capture the essence of cross-modal inputs. |
| Approach: | They propose to use vision-language models to generate conclusive phrases using multiple modalities of input information. |
| Outcome: | The proposed methods outperform existing methods on absence and unseen scenarios and overestimate model capability due to overlap in training tests. |
Scalable Vision Language Model Training via High Quality Data Curation (2025.acl-long)
Copied to clipboard
| Challenge: | SAIL-VL models achieve the highest average score in 18 widely used VLM benchmarks in our evaluation, with the 2B model takes the top position over VLMs of comparable sizes on OpenCompass 2024. |
| Approach: | They introduce an open-source vision language model (VLM) series that can be trained using high-quality data. |
| Outcome: | The proposed model achieves the highest average score in 18 widely used VLM benchmarks, with the 2B model taking the top position over VLMs of comparable sizes on OpenCompass 2024. |