Papers by Kaicheng Yang
CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of Modality (2020.acl-main)
Copied to clipboard
| Challenge: | Existing studies in multimodal sentiment analysis only use unified multimodal annotations, which do not reflect the independent sentiment of single modalities. |
| Approach: | They propose a Chinese single- and multi-modal sentiment analysis dataset with multimodal and independent unimodal annotations that can be used to study the interaction between modalities. |
| Outcome: | The proposed methods achieve state-of-the-art performance and learn more distinctive unimodal representations. |
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval (2025.emnlp-main)
Copied to clipboard
| Challenge: | a large-scale visionlanguage pre-training framework is limited by the scarcity of large-sized annotated vision-language data . noise-resistant data construction pipeline is needed to filter and caption web-sourced images . noisy text tokens can be a problem for fine-grained representation learning . |
| Approach: | They develop a noise-resistant data construction pipeline that leverages in-context learning capabilities of MLLMs to automatically filter and caption web-sourced images. |
| Outcome: | The proposed framework improves cross-modal alignment by masking noisy textual tokens based on the gradient-attention similarity score. |
RWKV-CLIP: A Robust Vision-Language Representation Learner (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using large image-text datasets, large-scale image-data sets have been used for visionlanguage pre-training. |
| Approach: | They propose a framework that leverages Large Language Models to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags. |
| Outcome: | The proposed framework can combine and refine information from web-based image-text pairs, synthetic captions, and detection tags. |