Papers by Cheongwoong Kang
Impact of Co-occurrence on Factual Knowledge of Large Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) often make factually incorrect responses despite their success in various applications. |
| Approach: | They propose to fine tune large language models to mitigate the co-occurrence bias by filtering out biased samples with high subject-object co-occurring counts. |
| Outcome: | The proposed model scales up to debiased datasets to mitigate the co-occurrence bias, but is not effective in recalling rare facts unseen during finetuning. |
When Format Changes Meaning: Investigating Semantic Inconsistency of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models are vulnerable to semantic inconsistency, a study finds . minor formatting variations result in divergent predictions for semantically equivalent inputs. |
| Approach: | They evaluate LLMs for semantic inconsistency and find they remain vulnerable . they propose to use mechanistic analysis to develop models that improve their reliability . |
| Outcome: | The proposed model is vulnerable to semantic inconsistency, the authors show . their model is brittle even in state-of-the-art models, they say . |
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation frameworks for large language models for domain specific tasks are coarse and do not provide a multidimensional evaluation of a model's ability to interpret domain specific data. |
| Approach: | They propose a diagnostic benchmark grounded in national qualification exams that exposes critical gaps across four dimensions: expert visual reasoning of charts, logical validity via expert-verified rationales, Korean-specific geo-cultural comprehension, and fine-grained domain analysis. |
| Outcome: | The proposed model outperforms global models in local contexts, demonstrating that parameter scaling alone cannot resolve cultural dependencies. |