Papers by Georgios Tzimiropoulos

6 papers
More Images, More Problems? A Controlled Analysis of VLM Failure Modes. (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of large vision language models lack a comprehensive analysis of their weaknesses and causes.
Approach: They propose a new benchmark to evaluate multi-image capabilities of Large Vision Language Models.
Outcome: The proposed model outperforms existing benchmarks on multi-image models.
Graph Guided Question Answer Generation for Procedural Question-Answering (2024.eacl-long)

Copied to clipboard

Challenge: a new method for question-answer generation from procedural text is sub-optimal for training QA models.
Approach: They propose a method for generating exhaustive and high-quality training data from procedural text . they use procedural data to represent each step and the overall flow of the procedure as graphs .
Outcome: The proposed method outperforms existing methods on task-specific question answering tasks.
Efficient Vision-Language pre-training via domain-specific learning for human activities (2024.emnlp-main)

Copied to clipboard

Challenge: Current vision-language models owe their success to large-scale pretraining on web-collected data.
Approach: They propose a domain-aligned pretraining strategy that aligns the downstream tasks to the downstream domain without additional data collection.
Outcome: The proposed method outperforms existing models on large-scale vision-language training datasets while preserving generalist knowledge.
MobileQuant: Mobile-friendly Quantization for On-device Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized language processing, but deployment on edge devices is costly in terms of memory, computation and energy.
Approach: They propose to reduce the number of bits used to represent weights and activations . they propose to use 8-bit activations to enable LLMs to fully exploit mobile-friendly hardware .
Outcome: The proposed method reduces the number of bits used to represent weights and activations . 8-bit activations are attractive for on-device deployment as they would exploit mobile-friendly hardware .
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions (2025.emnlp-main)

Copied to clipboard

Challenge: Contrastively trained Vision-Language Models exhibit shallow language understanding, manifesting bag-of-words behaviour.
Approach: They propose a vision-free, single-encoder retrieval pipeline to replace traditional text-to-image retrieval paradigm with structured image descriptions.
Outcome: The proposed approach reduces the modality gap and improves compositionality and performance on short and long caption queries.
A Simple Baseline for Knowledge-Based Visual Question Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies emphasize the importance of incorporating both explicit and implicit knowledge to answer questions requiring external knowledge.
Approach: They propose a pipeline that incorporates both explicit and implicit knowledge . their method is training-free and does not require access to external databases or APIs .
Outcome: The proposed method achieves state-of-the-art accuracy on OK-VQA and A-OK-VQ datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations