Papers by Seyoung Song
Are they lovers or friends? Evaluating LLMs’ Social Reasoning in English and Korean Dialogues (2026.acl-long)
Copied to clipboard
Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song, A. Seza Doğruöz, Alice Oh, Najoung Kim
| Challenge: | Existing studies on LLMs' ability to infer social relationships have limited results for Korean and English. |
| Approach: | They propose a social reasoning task based on a 1.1k-dialogue dataset in English and Korean sourced from movie scripts to evaluate LLMs' ability to infer the social relationships between speakers. |
| Outcome: | The proposed task evaluates the ability of LLMs to infer the social relationships between speakers in 1.1k-dialogue datasets in English and Korean. |
LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control (2026.acl-industry)
Copied to clipboard
| Challenge: | Using Large Language Models (LLMs) is challenging due to lack of domain-specific evaluation standards . current LLMs prioritize reasoning or knowledge over sociolinguistic nuances vital for automotive settings . |
| Approach: | They propose a framework for evaluation of Korean-language in-vehicle assistants . they propose to evaluate fine-grained Korean honorific control and safetycritical response behavior . |
| Outcome: | The proposed evaluation framework evaluates fine-grained honorific control, safetycritical response behavior, and task efficiency in deployment-aligned settings. |
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Evaluating text generation capabilities of large language models (LLMs) is challenging, especially for low-resource languages where methods for direct assessment are scarce. |
| Approach: | They propose a framework that transforms existing benchmarks into conversational tasks and measures LLMs’ accuracies on those tasks. |
| Outcome: | The proposed framework correlates strongly with established benchmarks while enabling standardized comparisons across languages and models. |