Papers by Guijin Son

11 papers
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: HAERAE-Vision benchmarks feature clear, explicit prompts but are often informal and underspecified . state-of-the-art models achieve under 50% on original queries, compared to GPT-5 and Gemini 2.5 Pro .
Approach: They propose a benchmark of 653 real-world visual questions from Korean online communities . they find that even state-of-the-art models achieve under 50% on original queries .
Outcome: HAERAE-Vision benchmarks from Korean online communities yield 1,306 query variants . state-of-the-art models achieve under 50% on original queries, compared with smaller models . authors show that query explicitation alone yields 8 to 22 point improvements .
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: a recent study has highlighted the fragility of Chain-of-Thought reasoning . a hypothesis suggests that effective communication is achieved by maintaining a stable flow of information.
Approach: They propose a framework to quantify uniformity of information flow at local and global levels . they propose entropy-based stepwise density metric to quantify this phenomenon .
Outcome: The proposed framework outperforms alternative signals as predictors of reasoning quality.
KMMLU: Measuring Massive Multitask Language Understanding in Korean (2025.naacl-long)

Copied to clipboard

Challenge: Recent models struggle to show performance over 60%, significantly below the pass mark of the source exams (80%), highlighting the room for improvement.
Approach: They propose to use Korean exams to collect 35,030 questions from an expert-level multiple choice model to capture linguistic and cultural aspects of the Korean language.
Outcome: The proposed benchmark is based on 35,030 questions from original Korean exams.
FINKRX: Establishing Best Practices for Korean Financial NLP (2025.acl-industry)

Copied to clipboard

Challenge: Existing tools to evaluate large language models in the financial domain are limited by the inherently closed nature of the financial industry.
Approach: They present the first open leaderboard for evaluating Korean large language models focused on finance.
Outcome: The proposed model is FINKRX, a fully open and transparent LLM built using these best practices.
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that pre-training compute can improve multilingual performance, but is it effective for test-time scaling?
Approach: They propose a multilingual math benchmark with competition-level problems in 55 languages . they propose "test-time scaling" which further lengthens the time it takes to scale .
Outcome: The proposed methods fail to generalize robustly across languages, with no improvements in variance or consistency.
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing evaluation tools rely on translations of English datasets or translation-specific benchmarks such as WMT 21 to assess large language models.
Approach: They propose a dataset curated to challenge models lacking Korean cultural and contextual depth.
Outcome: The HAE-RAE Bench challenges models lacking Korean cultural and contextual depth by highlighting their aptitude for recalling Korean-specific knowledge and cultural contexts.
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once? (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are typically trained to follow a single instruction per inference call.
Approach: They introduce a benchmark to evaluate Large language models' ability to follow one instruction per inference call.
Outcome: The proposed model reduces the total inference time by 1.46 times in average since it does not require multiple inference calls.
Controlling Language Confusion in Multilingual LLMs (2025.acl-srw)

Copied to clipboard

Challenge: Large language models suffer from language confusion, a phenomenon in which responses are partially or entirely generated in unintended languages.
Approach: They propose a supervised fine-tuning methodology which optimizes the likelihood of correct tokens without explicitly penalizing undesired outputs such as cross-lingual mixing.
Outcome: The proposed model suppresses language-confused generation while maintaining strong language consistency even under high decoding temperatures while preserving general QA performance.
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation (2025.findings-emnlp)

Copied to clipboard

Challenge: Using Korean expert-level benchmarks, Large Language Models can be developed in real-world scenarios.
Approach: They introduce two Korean expert-level benchmarks that reflect professional knowledge in Korea.
Outcome: The proposed benchmarks represent professional knowledge in Korea.
Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages? (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study focused on complex, high-level tasks, but LMentry is limited to English . a multilingual evaluation of large language models is needed to address this gap, authors say .
Approach: They propose a compact benchmark that enables systematic evaluation of large language models . they propose to use tasks that are trivial for humans but remain surprisingly difficult for LLMs .
Outcome: The proposed benchmark is limited to English, leaving its insights linguistically narrow.
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment.
Approach: They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation .
Outcome: The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations