Papers by Junsung Park

3 papers
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach (2024.acl-long)

Copied to clipboard

Challenge: primarily addressed in text-to-image retrieval task using dialogue-form context query . conventionally, text-based retrieval methods rely on initial text descriptions .
Approach: They propose a plug-based retrieval method that uses large language models as questioners to generate non-redundant questions about the attributes of the target image.
Outcome: The proposed method performs better than zero-shot and fine-tuned baselines in benchmarks.
Unleashing Multi-Hop Reasoning Potential in Large Language Models through Repetition of Misordered Context (2025.findings-naacl)

Copied to clipboard

Challenge: Multi-hop reasoning requires multi-step reasoning based on supporting documents within a given context.
Approach: They propose a method that prompts the model by repeatedly presenting the context.
Outcome: The proposed method improves the F1 score by 30%p on multi-hop QA tasks and increases accuracy by 70%p on a synthetic task.
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation (2025.findings-emnlp)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) have shown remarkable performance on diverse visual and linguistic tasks, yet they remain limited in their understanding of 3D spatial structures.
Approach: They propose a framework that injects human-inspired geometric cues into pretrained VLMs . they use sparse correspondences, relative depth relations and dense cost volumes .
Outcome: The proposed framework outperforms existing methods on vision-language reasoning and 3D perception benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations