Papers by Xinyuan Cai
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding (2026.findings-acl)
Copied to clipboard
| Challenge: | Currently, vision-Language Models are optimized for direct visual question-answering tasks. |
| Approach: | They propose a visual-language-based VLM that prioritizes reasoning within the perception process. |
| Outcome: | The proposed model outperforms existing models and domain-specific open-source models in the chemical domain. |
Dash-M5H: An Interactive Dashboard for Multi-Modal, Multi-Model Mental Health Assessment (2026.acl-demo)
Copied to clipboard
Raymond Alavo, Xinyuan Zhang, Gemza Ademaj, Junhui Cai, Hyeokhyen Kwon, Robert Cotes, Gari D. Clifford, Ahmed Abbasi
| Challenge: | Dash-M5H integrates transcript text, audio, and facial behavior with a clinically grounded VLM prediction pipeline that produces DSM-5-aligned depression predictions. |
| Approach: | They propose a dashboard that integrates multimodal behavioral data with multi-model signal outputs of recorded clinical interviews. |
| Outcome: | Dash-M5H is an interactive dashboard for *multi-modal, multi-model mental health assessment that integrates transcript text, audio, and facial behavior with a clinically grounded VLM prediction pipeline. |