Papers by HaiXiang Zhu
Read Before Grounding: Scene Knowledge Visual Grounding via Multi-step Parsing (2025.coling-main)
Copied to clipboard
| Challenge: | Existing VG datasets use simple textual descriptions with limited attribute and spatial information between images and text. |
| Approach: | They propose a method that transforms visual knowledge into concise, information-dense visual descriptions. |
| Outcome: | The proposed method significantly improves performance of multimodal grounding models. |