Papers by Kwonjoon Lee

5 papers
UQ-Merge: Uncertainty Guided Multimodal Large Language Model Merging (2025.findings-acl)

Copied to clipboard

Challenge: Existing models merging methods often lead to suboptimal performance due to harmful models . et al., 2018; 59: 59-64.
Approach: They propose an uncertainty-guided MLLM merging algorithm that integrates models into a single MLML.
Outcome: The proposed algorithm improves on held-in and held-out vision-language benchmarks.
Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Model based multi-agent systems (MAS) excel at collaborative problem solving but remain brittle to cascading errors.
Approach: They propose a metacognitive framework that enables step-level error detection and self-correction in Large Language Model based multi-agent systems (MAS) .
Outcome: The proposed framework outperforms baselines on the Who When benchmark and delivers consistent gains on AgentErrorBench.
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for visual commonsense reasoning (VCR) use pre-trained large language models and pre-training visionlanguage models.
Approach: They propose a collaborative approach where pre-trained LLMs serve as problem classifiers to analyze problem category and either use VLMs to answer directly or actively instruct LLM to gather relevant visual elements to support potential commonsense inferences.
Outcome: The proposed approach outperforms all other methods without in-domain fine-tuning on two VCR benchmark datasets.
Task-Aware Resolution Optimization for Visual Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing visual large language models pre-assume a fixed resolution for downstream tasks, leading to sub-optimal performance.
Approach: They propose a formula to determine the optimal resolution for a given vision-language task . they then propose 'parameter-efficient' fine-tuning technique to extend the visual input resolution .
Outcome: The proposed method is based on rigorous experiments on vision-language tasks.
Can Hallucination Correction Improve Video-Language Alignment? (2025.findings-acl)

Copied to clipboard

Challenge: Existing work on hallucination correction for large vision-language models focuses on mitigating hallucisations, but a new approach is needed to improve video-language alignment.
Approach: They propose a self-training framework learning to correct hallucinations in descriptions that do not align with the video content.
Outcome: The proposed framework improves video-language alignment by identifying and correcting inconsistencies in descriptions that do not align with the video content.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations