Papers by Guijin Son
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models (2026.findings-acl)
Copied to clipboard
Dasol Choi, Guijin Son, Hanwool Lee, Minhyuk Kim, Hyunwoo Ko, Teabin Lim, Eungyeol Ahn, Jungwhan Kim, Seunghyeok Hong, Youngsook Song
| Challenge: | HAERAE-Vision benchmarks feature clear, explicit prompts but are often informal and underspecified . state-of-the-art models achieve under 50% on original queries, compared to GPT-5 and Gemini 2.5 Pro . |
| Approach: | They propose a benchmark of 653 real-world visual questions from Korean online communities . they find that even state-of-the-art models achieve under 50% on original queries . |
| Outcome: | HAERAE-Vision benchmarks from Korean online communities yield 1,306 query variants . state-of-the-art models achieve under 50% on original queries, compared with smaller models . authors show that query explicitation alone yields 8 to 22 point improvements . |
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | a recent study has highlighted the fragility of Chain-of-Thought reasoning . a hypothesis suggests that effective communication is achieved by maintaining a stable flow of information. |
| Approach: | They propose a framework to quantify uniformity of information flow at local and global levels . they propose entropy-based stepwise density metric to quantify this phenomenon . |
| Outcome: | The proposed framework outperforms alternative signals as predictors of reasoning quality. |
KMMLU: Measuring Massive Multitask Language Understanding in Korean (2025.naacl-long)
Copied to clipboard
Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, Stella Biderman
| Challenge: | Recent models struggle to show performance over 60%, significantly below the pass mark of the source exams (80%), highlighting the room for improvement. |
| Approach: | They propose to use Korean exams to collect 35,030 questions from an expert-level multiple choice model to capture linguistic and cultural aspects of the Korean language. |
| Outcome: | The proposed benchmark is based on 35,030 questions from original Korean exams. |
FINKRX: Establishing Best Practices for Korean Financial NLP (2025.acl-industry)
Copied to clipboard
| Challenge: | Existing tools to evaluate large language models in the financial domain are limited by the inherently closed nature of the financial industry. |
| Approach: | They present the first open leaderboard for evaluating Korean large language models focused on finance. |
| Outcome: | The proposed model is FINKRX, a fully open and transparent LLM built using these best practices. |
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning (2025.acl-long)
Copied to clipboard
| Challenge: | Recent studies show that pre-training compute can improve multilingual performance, but is it effective for test-time scaling? |
| Approach: | They propose a multilingual math benchmark with competition-level problems in 55 languages . they propose "test-time scaling" which further lengthens the time it takes to scale . |
| Outcome: | The proposed methods fail to generalize robustly across languages, with no improvements in variance or consistency. |
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models (2024.lrec-main)
Copied to clipboard
Guijin Son, Hanwool Lee, Suwan Kim, Huiseo Kim, Jae cheol Lee, Je Won Yeom, Jihyu Jung, Jung woo Kim, Songseong Kim
| Challenge: | Existing evaluation tools rely on translations of English datasets or translation-specific benchmarks such as WMT 21 to assess large language models. |
| Approach: | They propose a dataset curated to challenge models lacking Korean cultural and contextual depth. |
| Outcome: | The HAE-RAE Bench challenges models lacking Korean cultural and contextual depth by highlighting their aptitude for recalling Korean-specific knowledge and cultural contexts. |
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once? (2024.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are typically trained to follow a single instruction per inference call. |
| Approach: | They introduce a benchmark to evaluate Large language models' ability to follow one instruction per inference call. |
| Outcome: | The proposed model reduces the total inference time by 1.46 times in average since it does not require multiple inference calls. |
Controlling Language Confusion in Multilingual LLMs (2025.acl-srw)
Copied to clipboard
| Challenge: | Large language models suffer from language confusion, a phenomenon in which responses are partially or entirely generated in unintended languages. |
| Approach: | They propose a supervised fine-tuning methodology which optimizes the likelihood of correct tokens without explicitly penalizing undesired outputs such as cross-lingual mixing. |
| Outcome: | The proposed model suppresses language-confused generation while maintaining strong language consistency even under high decoding temperatures while preserving general QA performance. |
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Using Korean expert-level benchmarks, Large Language Models can be developed in real-world scenarios. |
| Approach: | They introduce two Korean expert-level benchmarks that reflect professional knowledge in Korea. |
| Outcome: | The proposed benchmarks represent professional knowledge in Korea. |
Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages? (2025.emnlp-main)
Copied to clipboard
Luca Moroni, Javier Aula-Blasco, Simone Conia, Irene Baucells, Naiara Perez, Silvia Paniagua Suárez, Anna Sallés, Malte Ostendorff, Júlia Falcão, Guijin Son, Aitor Gonzalez-Agirre, Roberto Navigli, Marta Villegas
| Challenge: | a recent study focused on complex, high-level tasks, but LMentry is limited to English . a multilingual evaluation of large language models is needed to address this gap, authors say . |
| Approach: | They propose a compact benchmark that enables systematic evaluation of large language models . they propose to use tasks that are trivial for humans but remain surprisingly difficult for LLMs . |
| Outcome: | The proposed benchmark is limited to English, leaving its insights linguistically narrow. |
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)
Copied to clipboard
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo
| Challenge: | a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment. |
| Approach: | They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation . |
| Outcome: | The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks. |