Papers by Eunsu Kim

10 papers
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean (2024.lrec-main)

Copied to clipboard

Challenge: Existing benchmark datasets for Korean cultural and linguistic knowledge are derived from the English counterparts through translation, so they overlook cultural contexts.
Approach: They propose to use Korean cultural and linguistic intelligence to assess Korean model performance by providing fine-grained annotations of cultural and cultural knowledge.
Outcome: The proposed dataset includes 1,995 QA pairs and is based on 1,992 Korean exams and textbooks.
Are they lovers or friends? Evaluating LLMs’ Social Reasoning in English and Korean Dialogues (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on LLMs' ability to infer social relationships have limited results for Korean and English.
Approach: They propose a social reasoning task based on a 1.1k-dialogue dataset in English and Korean sourced from movie scripts to evaluate LLMs' ability to infer the social relationships between speakers.
Outcome: The proposed task evaluates the ability of LLMs to infer the social relationships between speakers in 1.1k-dialogue datasets in English and Korean.
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Recent work on LLM-as-a-Judge has reported higher correlations with human judgments due to its static nature.
Approach: They propose a framework that leverages multi-turn interactions where the LLM interviewer actively provides feedback on responses and poses follow-up questions to the evaluated LLM.
Outcome: The proposed framework evaluates six models on reasoning, factuality and instruction-following tasks.
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation methods struggle to capture subtle inconsistencies in large language models.
Approach: They propose an atomic-level evaluation framework that quantifies persona fidelity at a finer granularity.
Outcome: The proposed framework detects inconsistencies that prior evaluation methods overlook . it captures subtle deviations that real users would encounter .
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs’ Contextual Sensitivity (2026.findings-acl)

Copied to clipboard

Challenge: a new benchmark measures the contextual sensitivity of large language models in role conflict scenarios . role conflicts are social dilemmas where multiple roles cannot be fulfilled simultaneously . authors: models are forced to arbitrate between dynamic contextual cues and learned preferences .
Approach: They propose a benchmark to measure the contextual sensitivity of large language models in role conflict scenarios.
Outcome: The proposed benchmark measures the contextual sensitivity of large language models in role conflict scenarios.
Diffusion Models Through a Global Lens: Are They Culturally Inclusive? (2025.acl-long)

Copied to clipboard

Challenge: Text-to-image diffusion models have produced compelling, detailed images from text prompts, but their ability to accurately represent cultural nuances remains an open question.
Approach: They propose a benchmark to evaluate whether diffusion models can generate culturally specific images spanning ten countries.
Outcome: The proposed model fails to generate culturally specific images spanning ten countries . it shows significant disparities in cultural relevance, description fidelity, and realism compared to real-world reference images.
Uncovering Factor-Level Preference to Improve Human-Model Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models exhibit tendencies that diverge from human preferences, such as favoring certain writing styles or producing overly verbose outputs.
Approach: They propose a framework to uncover and measure factor-level preference alignment of humans and large language models (LLMs)
Outcome: The proposed framework uncovers and measures factor-level preference alignment of humans and large language models.
The Generative AI Paradox in Evaluation: “What It Can Solve, It May Not Evaluate” (2024.eacl-srw)

Copied to clipboard

Challenge: Existing studies on using Large Language Models for model evaluation have focused on using LLMs for reference-free evaluation to meet the needs of long-form text evaluation.
Approach: They propose to use Large Language Models (LLMs) for generation tasks to evaluate models.
Outcome: The proposed model evaluations show that LLMs are less faithful to evaluation tasks than open-source models.
LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control (2026.acl-industry)

Copied to clipboard

Challenge: Using Large Language Models (LLMs) is challenging due to lack of domain-specific evaluation standards . current LLMs prioritize reasoning or knowledge over sociolinguistic nuances vital for automotive settings .
Approach: They propose a framework for evaluation of Korean-language in-vehicle assistants . they propose to evaluate fine-grained Korean honorific control and safetycritical response behavior .
Outcome: The proposed evaluation framework evaluates fine-grained honorific control, safetycritical response behavior, and task efficiency in deployment-aligned settings.
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language (2025.findings-emnlp)

Copied to clipboard

Challenge: Evaluating text generation capabilities of large language models (LLMs) is challenging, especially for low-resource languages where methods for direct assessment are scarce.
Approach: They propose a framework that transforms existing benchmarks into conversational tasks and measures LLMs’ accuracies on those tasks.
Outcome: The proposed framework correlates strongly with established benchmarks while enabling standardized comparisons across languages and models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations