Papers by Michael Ogezi

2 papers
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)

Copied to clipboard

Challenge: Vision-language models struggle with spatial reasoning, a skill that humans excel at.
Approach: They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics.
Outcome: The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks.
Semantically-Prompted Language Models Improve Visual Descriptions (2024.findings-naacl)

Copied to clipboard

Challenge: Language-vision models have made significant progress in zeroshot vision tasks, but lack expressive visual descriptions.
Approach: They propose a new method for generating visual descriptions with pre-trained language models and semantic knowledge bases.
Outcome: The proposed method improves visual descriptions and achieves strong results on image-classification datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations