Papers by Kyoungson Jhang
PseudoGD: Enhancing Spatial Reasoning in Vision-Language Models through Pseudo Geometric Knowledge Distillation (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent Large Vision-Language Models (LVLMs) have shown remarkable success in general semantic understanding, but struggle with 3D spatial reasoning tasks. |
| Approach: | They propose a framework to help vision encoders internalize 3D geometric information using only standard 2D images. |
| Outcome: | The proposed framework achieves State-of-the-Art (SOTA) performance across various model architectures. |