Papers by Dongnan Yang
RWKV-CLIP: A Robust Vision-Language Representation Learner (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using large image-text datasets, large-scale image-data sets have been used for visionlanguage pre-training. |
| Approach: | They propose a framework that leverages Large Language Models to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags. |
| Outcome: | The proposed framework can combine and refine information from web-based image-text pairs, synthetic captions, and detection tags. |
Dual-Path Dynamic Fusion with Learnable Query for Multimodal Sentiment Analysis (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for multimodal sentiment analysis struggle with global and fine-grained contributions and over-reliance on text. |
| Approach: | They propose a multimodal sentiment analysis architecture that processes inputs through two complementary paths: global and local. |
| Outcome: | The proposed architecture achieves state-of-the-art in fine-grained sentiment prediction on the CMU-MOSI and CMU MOSEI benchmarks. |