Papers by Rafid Mahmood
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Despite recent advances in visual language models, their ability to quantitatively reason about object sizes and distances remains underexplored. |
| Approach: | They propose a manually annotated benchmark of 241 questions designed for quantitative spatial reasoning and a zero-shot prompting technique that encourages VLMs to use reference objects as visual cues. |
| Outcome: | The proposed technique improves the performance of the top-performing VLMs by 19 points when a reasoning path using a reference object emerges naturally in the response. |