Papers by Soyeon Kim
DeFrame: Debiasing Large Language Models Against Framing Effects (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing debiasing methods improve overall fairness, but fail to reduce framing-induced disparities. |
| Approach: | They propose a framing-aware debiasing method that encourages LLMs to be more consistent across frams. |
| Outcome: | The proposed method reduces overall bias and improves robustness against framing disparities, enabling LLMs to produce fairer and more consistent responses. |
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Using Korean expert-level benchmarks, Large Language Models can be developed in real-world scenarios. |
| Approach: | They introduce two Korean expert-level benchmarks that reflect professional knowledge in Korea. |
| Outcome: | The proposed benchmarks represent professional knowledge in Korea. |
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation frameworks for large language models for domain specific tasks are coarse and do not provide a multidimensional evaluation of a model's ability to interpret domain specific data. |
| Approach: | They propose a diagnostic benchmark grounded in national qualification exams that exposes critical gaps across four dimensions: expert visual reasoning of charts, logical validity via expert-verified rationales, Korean-specific geo-cultural comprehension, and fine-grained domain analysis. |
| Outcome: | The proposed model outperforms global models in local contexts, demonstrating that parameter scaling alone cannot resolve cultural dependencies. |
Hospitality-VQA: Decision-Oriented Informativeness Evaluation for Vision–Language Models (2026.eacl-srw)
Copied to clipboard
Jeongwoo Lee, Baek Duhyeong, Eungyeol Han, Soyeon Shin, Gukin Han, Seungduk Kim, Jaehyun Jeon, Taewoo Jeong
| Challenge: | Existing VQA benchmarks focus on factual correctness but rarely capture what information users actually find useful. |
| Approach: | They propose a framework to quantify how much information an image–question pair provides . they conduct experiments with several state-of-the-art VLMs to determine their reliability . |
| Outcome: | The proposed framework quantifies how much information an image–question pair provides in hospitality contexts. |
Connecting the Knowledge Dots: Retrieval-augmented Knowledge Connection for Commonsense Reasoning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that large language models exhibit a limited understanding of commonsense reasoning due to the necessity of implicit knowledge that is rarely expressed in text. |
| Approach: | They propose a retrieval-augmented knowledge connection framework that transforms indirectly relevant documents into a direct explanation to answer a given question. |
| Outcome: | The proposed framework outperforms state-of-the-art (SOTA) benchmarks and achieves +2.0% and +4.6% average accuracy on in-domain (ID) and out-of domain (OOD) benchmark. |