Papers by Jaehyun Jeon

4 papers
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you! (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models lack this active understanding capacity, limiting their applicability in real-world scenarios.
Approach: They propose a benchmark to assess the impact of multimodal inputs on lexical ambiguities.
Outcome: The proposed benchmark assesses the impact of multimodal inputs on lexical ambiguities.
Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding (2026.acl-long)

Copied to clipboard

Challenge: Recent studies focus on surface-level features, overlooking how design choices influence user behavior at scale.
Approach: They propose a benchmark for multimodal understanding of how UI/UX design affects user behavior built on 300 real-world UI image pairs from industry A/B tests.
Outcome: The proposed benchmarks show that models exhibit limited understanding of the behavioral impact of UI/UX design.
Hospitality-VQA: Decision-Oriented Informativeness Evaluation for Vision–Language Models (2026.eacl-srw)

Copied to clipboard

Challenge: Existing VQA benchmarks focus on factual correctness but rarely capture what information users actually find useful.
Approach: They propose a framework to quantify how much information an image–question pair provides . they conduct experiments with several state-of-the-art VLMs to determine their reliability .
Outcome: The proposed framework quantifies how much information an image–question pair provides in hospitality contexts.
Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal large language models struggle when faced with unseen domains or languages.
Approach: They propose a framework that leverages the broad knowledge of an MLLM to generate cross-modal pre-questions (preQs) before retrieval.
Outcome: Experiments show that PREMIR outperforms existing retrievers on out-of-distribution benchmarks, including closed-domain and multilingual settings, outperforming strong baselines across all metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations