Papers by Marjan Alirezaie
CLEVR-POC: Reasoning-Intensive Visual Question Answering in Partially Observable Environments (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing knowledge for reasoning about partially observed scenes is limited . lucian et al. show that pre-trained vision language models are not adequate for VQA . |
| Approach: | They propose a benchmark for reasoning-intensive visual question answering . they use logical constraints to leverage knowledge to generate plausible answers . |
| Outcome: | The proposed model performs better than pre-trained models on CLEVR-POC . the proposed model is based on a neuro-symbolic model with a visual perception network and a formal logical reasoner . |