Papers by Eunsu Kim
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing benchmark datasets for Korean cultural and linguistic knowledge are derived from the English counterparts through translation, so they overlook cultural contexts. |
| Approach: | They propose to use Korean cultural and linguistic intelligence to assess Korean model performance by providing fine-grained annotations of cultural and cultural knowledge. |
| Outcome: | The proposed dataset includes 1,995 QA pairs and is based on 1,992 Korean exams and textbooks. |
Are they lovers or friends? Evaluating LLMs’ Social Reasoning in English and Korean Dialogues (2026.acl-long)
Copied to clipboard
Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song, A. Seza Doğruöz, Alice Oh, Najoung Kim
| Challenge: | Existing studies on LLMs' ability to infer social relationships have limited results for Korean and English. |
| Approach: | They propose a social reasoning task based on a 1.1k-dialogue dataset in English and Korean sourced from movie scripts to evaluate LLMs' ability to infer the social relationships between speakers. |
| Outcome: | The proposed task evaluates the ability of LLMs to infer the social relationships between speakers in 1.1k-dialogue datasets in English and Korean. |
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent work on LLM-as-a-Judge has reported higher correlations with human judgments due to its static nature. |
| Approach: | They propose a framework that leverages multi-turn interactions where the LLM interviewer actively provides feedback on responses and poses follow-up questions to the evaluated LLM. |
| Outcome: | The proposed framework evaluates six models on reasoning, factuality and instruction-following tasks. |
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation methods struggle to capture subtle inconsistencies in large language models. |
| Approach: | They propose an atomic-level evaluation framework that quantifies persona fidelity at a finer granularity. |
| Outcome: | The proposed framework detects inconsistencies that prior evaluation methods overlook . it captures subtle deviations that real users would encounter . |
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs’ Contextual Sensitivity (2026.findings-acl)
Copied to clipboard
| Challenge: | a new benchmark measures the contextual sensitivity of large language models in role conflict scenarios . role conflicts are social dilemmas where multiple roles cannot be fulfilled simultaneously . authors: models are forced to arbitrate between dynamic contextual cues and learned preferences . |
| Approach: | They propose a benchmark to measure the contextual sensitivity of large language models in role conflict scenarios. |
| Outcome: | The proposed benchmark measures the contextual sensitivity of large language models in role conflict scenarios. |
Diffusion Models Through a Global Lens: Are They Culturally Inclusive? (2025.acl-long)
Copied to clipboard
Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad, Eunsu Kim, Junyeong Park, James Thorne, Alice Oh
| Challenge: | Text-to-image diffusion models have produced compelling, detailed images from text prompts, but their ability to accurately represent cultural nuances remains an open question. |
| Approach: | They propose a benchmark to evaluate whether diffusion models can generate culturally specific images spanning ten countries. |
| Outcome: | The proposed model fails to generate culturally specific images spanning ten countries . it shows significant disparities in cultural relevance, description fidelity, and realism compared to real-world reference images. |
Uncovering Factor-Level Preference to Improve Human-Model Alignment (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models exhibit tendencies that diverge from human preferences, such as favoring certain writing styles or producing overly verbose outputs. |
| Approach: | They propose a framework to uncover and measure factor-level preference alignment of humans and large language models (LLMs) |
| Outcome: | The proposed framework uncovers and measures factor-level preference alignment of humans and large language models. |
The Generative AI Paradox in Evaluation: “What It Can Solve, It May Not Evaluate” (2024.eacl-srw)
Copied to clipboard
| Challenge: | Existing studies on using Large Language Models for model evaluation have focused on using LLMs for reference-free evaluation to meet the needs of long-form text evaluation. |
| Approach: | They propose to use Large Language Models (LLMs) for generation tasks to evaluate models. |
| Outcome: | The proposed model evaluations show that LLMs are less faithful to evaluation tasks than open-source models. |
LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control (2026.acl-industry)
Copied to clipboard
| Challenge: | Using Large Language Models (LLMs) is challenging due to lack of domain-specific evaluation standards . current LLMs prioritize reasoning or knowledge over sociolinguistic nuances vital for automotive settings . |
| Approach: | They propose a framework for evaluation of Korean-language in-vehicle assistants . they propose to evaluate fine-grained Korean honorific control and safetycritical response behavior . |
| Outcome: | The proposed evaluation framework evaluates fine-grained honorific control, safetycritical response behavior, and task efficiency in deployment-aligned settings. |
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Evaluating text generation capabilities of large language models (LLMs) is challenging, especially for low-resource languages where methods for direct assessment are scarce. |
| Approach: | They propose a framework that transforms existing benchmarks into conversational tasks and measures LLMs’ accuracies on those tasks. |
| Outcome: | The proposed framework correlates strongly with established benchmarks while enabling standardized comparisons across languages and models. |