Papers by Amir Bar

3 papers
Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models perform well on visual question answering tasks, but it remains unclear whether their reasoning relies more on memorized world knowledge or on visual information present in the input image.
Approach: They propose a dataset of visual-realistic counterfactuals that put world knowledge priors into conflict with visual input.
Outcome: The proposed dataset puts world knowledge priors into conflict with visual input . it shows that model predictions shift toward visual evidence in mid-to-late layers .
DiaSet: An Annotated Dataset of Arabic Conversations (2024.lrec-main)

Copied to clipboard

Challenge: DiaSet is a dataset of dialectical Arabic speech manually transcribed and annotated for two downstream tasks.
Approach: They propose to manually transcribe and annotate Arabic speech for sentiment analysis and named entity recognition.
Outcome: The proposed dataset encapsulates the Palestine dialect, predominantly spoken in Palestine, Israel, and Jordan.
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models struggle with visual reasoning, despite strong performance on vision-language tasks.
Approach: They propose a visually cued chain-of-thought prompting that enhances multi-step mathematical reasoning by explicitly referencing visual annotations in diagrams.
Outcome: The proposed model improves GPT-4o's accuracy on an irregular polygon side-counting task from 7% to 93%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations