Papers by Ran Jiao

2 papers
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Multi-modal keyphrase prediction (MMKP) aims to produce concise, informative phrases that capture the essence of cross-modal inputs.
Approach: They propose to use vision-language models to generate conclusive phrases using multiple modalities of input information.
Outcome: The proposed methods outperform existing methods on absence and unseen scenarios and overestimate model capability due to overlap in training tests.
Scalable Vision Language Model Training via High Quality Data Curation (2025.acl-long)

Copied to clipboard

Challenge: SAIL-VL models achieve the highest average score in 18 widely used VLM benchmarks in our evaluation, with the 2B model takes the top position over VLMs of comparable sizes on OpenCompass 2024.
Approach: They introduce an open-source vision language model (VLM) series that can be trained using high-quality data.
Outcome: The proposed model achieves the highest average score in 18 widely used VLM benchmarks, with the 2B model taking the top position over VLMs of comparable sizes on OpenCompass 2024.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations