Papers by Michael Ogezi
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)
Copied to clipboard
| Challenge: | Vision-language models struggle with spatial reasoning, a skill that humans excel at. |
| Approach: | They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics. |
| Outcome: | The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks. |
Semantically-Prompted Language Models Improve Visual Descriptions (2024.findings-naacl)
Copied to clipboard
| Challenge: | Language-vision models have made significant progress in zeroshot vision tasks, but lack expressive visual descriptions. |
| Approach: | They propose a new method for generating visual descriptions with pre-trained language models and semantic knowledge bases. |
| Outcome: | The proposed method improves visual descriptions and achieves strong results on image-classification datasets. |