Papers by Matthew Olson

1 papers
Why do LLaVA Vision-Language Models Reply to Images in English? (2024.findings-emnlp)

Copied to clipboard

Challenge: Including an image in a multimodal query significantly increases the likelihood of the model returning an English response regardless of the language of the query.
Approach: They propose a two-pronged approach that combines extensive ablation of the design space with a mechanistic analysis of the models’ internal representations of image and text inputs.
Outcome: The proposed approach reduces the multilingual error by switching the language backbone for a bilingual language model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations