Papers by Menglan Chen
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization (2026.findings-acl)
Copied to clipboard
| Challenge: | integrating vision and language models with safety standards is essential to mitigate multimodal complexity . integrating visual inputs with vision and text unveils subtle threats beyond the reach of conventional safeguards . |
| Approach: | They propose a framework that combines vision and language to provide a multimodal reasoning-driven prompt rewriting. |
| Outcome: | The proposed framework outperforms baseline models on five benchmarks with six VLMs. |