Papers by Can Qin

2 papers
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Vision Language Models struggle with visual arithmetic, seemingly simple tasks like object counting or length comparison, which are essential for relevant complex tasks like chart understanding and geometric reasoning.
Approach: They propose a novel post-training strategy inspired by Piaget’s theory of cognitive development that trains VLMs to recognize invariant properties under visual transformations.
Outcome: The proposed approach outperforms supervised fine-tuning methods while requiring 60% less training data.
Self-Training Large Language and Vision Assistant for Medical Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for collecting medical data are expensive and time-consuming.
Approach: They propose a method to train a large-scale LVLM capable of auto-generating medical visual instruction data to improve data efficiency.
Outcome: The proposed method shows that it performs well across three major visual question answering (VQA) benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations