Papers by Joonhyung Park

4 papers
Bringing Real-World Relations into Video Generation with Graph-Structured Knowledge (2026.acl-long)

Copied to clipboard

Challenge: Existing text-to-video models struggle to accurately simulate real-world physics and dynamic entity interactions.
Approach: They propose a framework that integrates graph-structured temporal knowledge into video latent diffusion models to enhance compositional generation and interaction fidelity.
Outcome: The proposed framework enhances compositional generation and interaction fidelity by integrating graph-structured temporal knowledge into video latent diffusion models.
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding (2025.findings-acl)

Copied to clipboard

Challenge: Existing vision-only GUI agents ground elements from large and cluttered screenshots, requiring them to process substantial irrelevant information that compromises their accuracy.
Approach: They propose a visual agent model for GUI automation that leverages zoomed-in region proposals for precise element localization.
Outcome: The proposed approach improves state-of-the-art grounding accuracy by 13% across diverse GUI platforms on the GUI grounding benchmarks ScreenSpot and AgentStudio.
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens (2025.emnlp-main)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) generate contextually relevant responses by jointly interpreting visual and textual inputs.
Approach: They propose a method to classify whether an input token is visually grounded by reinterpreting question prompts or replacing the detected absent tokens during generation.
Outcome: The proposed method mitigates the models’ tendency to falsely presume the visual presence of text input and its generality across various LVLMs.
Language-Grounded Multi-Domain Image Translation via Semantic Difference Guidance (2026.eacl-long)

Copied to clipboard

Challenge: Existing methods for image-to-image translation lack structural integrity and attribute-specific control . Existing approaches lack semantics and provide fine-grained, attribute-based control compared to GAN-based methods .
Approach: They propose a language-grounded attribute-controllable translation framework that grounds semantic differences into corresponding visual transformations while preserving unrelated structural and semantic content.
Outcome: Experiments on CelebA(Dialog) and BDD100K show that LACE achieves high visual fidelity, structural preservation, and interpretable domain-specific control, surpassing baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations