Papers by Shiming Xiang
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for Knowledge-Based Visual Question Answering rely on images as the retrieval key, and often overlook or misplace the role of Vision-Language Models (VLMs) |
| Approach: | They propose a multi-modal RAG framework that assigns VLMs two specialized agents: a Refiner and an Inspector. |
| Outcome: | Experiments on EVQA, InfoSeek, and M2KR show that the proposed framework achieves state-of-the-art performance with significant improvements in both retrieval accuracy and answer quality. |
Faster and Better LLMs via Latency-Aware Test-Time Scaling (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing research has overlooked the efficiency of TTS from a latency-sensitive perspective. |
| Approach: | They propose two approaches to achieve latency-optimal TTS by branch-wise parallelism and sequence-wise parallelism. |
| Outcome: | The proposed approach achieves latency-optimal TTS for large models . branch-wise parallelism and sequence-wise parallelism are key approaches . |