Papers by Jaehyung Seo
A Dog Is Passing Over The Jet? A Text-Generation Dataset for Korean Commonsense Reasoning and Evaluation (2022.findings-naacl)
Copied to clipboard
Jaehyung Seo, Seounghoon Lee, Chanjun Park, Yoonna Jang, Hyeonseok Moon, Sugyeong Eo, Seonmin Koo, Heuiseok Lim
| Challenge: | Korean pretrained language models struggle to generate short sentences with a given condition based on compositionality and commonsense reasoning. |
| Approach: | They propose a Korean text-generation dataset for Korean generative commonsense reasoning and language model evaluation using a semi-automatic dataset construction approach. |
| Outcome: | The proposed dataset is available at http://aihub.or.kr/opendata/korea-university. |
Detecting Critical Errors Considering Cross-Cultural Factors in English-Korean Translation (2024.lrec-main)
Copied to clipboard
Sugyeong Eo, Jungwoo Lim, Chanjun Park, DaHyun Jung, Seonmin Koo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim
| Challenge: | Recent machine translation systems overcome language barriers for a wide range of users, yet they carry the risk of catastrophic meaning deviations. |
| Approach: | They introduce a culture-aware "Politeness" type for detecting critical translation errors . they also provide multiclass labels for critical error detection and critical error type classification . |
| Outcome: | Empirical results show that the proposed method outperforms baselines in both tasks. |
Call for Rigor in Reporting Quality of Instruction Tuning Data (2025.acl-short)
Copied to clipboard
| Challenge: | Instruction tuning is crucial for adapting large language models (LLMs) to user intentions. |
| Approach: | They propose to use hyperparameters for training models that are often selected arbitrarily without adequate justification to make arbitrary conclusions. |
| Outcome: | The results show that arbitrary hyperparameter decisions can make any arbitrary conclusion. |
LimaCost: Data Valuation for Instruction Tuning of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Instruction tuning is an effective approach for aligning large language models with human intentions. |
| Approach: | They propose a data quality measure that exhibits a strong correlation with model performance. |
| Outcome: | The proposed measure exhibits a strong correlation with model performance. |
Length-aware Byte Pair Encoding for Mitigating Over-segmentation in Korean Machine Translation (2024.findings-acl)
Copied to clipboard
Jungseob Lee, Hyeonseok Moon, Seungjun Lee, Chanjun Park, Sugyeong Eo, Hyunwoong Ko, Jaehyung Seo, Seungyoon Lee, Heuiseok Lim
| Challenge: | Byte Pair Encoding (BPE) is an effective approach in machine translation across several languages, but it is prone to over-segmentation in Korean, an agglutinative and morphologically rich language. |
| Approach: | They propose a new method that incorporates long words into the Korean vocabulary by strategically preserving morphological information and reducing semantic confusion. |
| Outcome: | The proposed method outperforms BPE and surpasses state-of-the-art morpheme-aware tokenization methods. |
CReTIHC: Designing Causal Reasoning Tasks about Temporal Interventions and Hallucinated Confoundings (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive capabilities in natural language processing, but their ability to establish causal relationships remains challenging. |
| Approach: | They propose a novel dataset to test and enhance the causal reasoning abilities of large language models (LLMs) by integrating elements of verbal hallucinations and temporal interventions into existing causal inference datasets. |
| Outcome: | The proposed dataset is designed to test and enhance the causal reasoning abilities of large language models. |
PEEP-Talk: A Situational Dialogue-based Chatbot for English Education (2023.acl-demo)
Copied to clipboard
Seungjun Lee, Yoonna Jang, Chanjun Park, Jungseob Lee, Jaehyung Seo, Hyeonseok Moon, Sugyeong Eo, Seounghoon Lee, Bernardo Yahya, Heuiseok Lim
| Challenge: | Existing chatbots lack realistic practice scenarios for English learners . existing platforms employ hand-crafted and patternmatching rules, limiting communication ability and responding appropriately to out-of-situation utterances. |
| Approach: | They propose a real-world situational dialogue-based chatbot for English education . it generates appropriate responses in various real-life situations while providing accurate feedback . |
| Outcome: | The proposed chatbot generates appropriate responses in various real-life situations while providing accurate feedback to learners. |
CoME: An Unlearning-based Approach to Conflict-free Model Editing (2025.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) often retain outdated or incorrect information from pre-training, which undermines their reliability. |
| Approach: | They propose a conflict-free model editing framework that selectively removes outdated knowledge from LLMs to improve their accuracy and reliability. |
| Outcome: | The proposed framework improves both editing accuracy and model reliability when applied to existing editing methods. |
KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Language models are striving to grasp commonsense reasoning, but they are lacking in Korean commons- ense benchmarks. |
| Approach: | They present a fine-grained benchmark dataset focused on Korean commonsense reasoning that includes multiple-choice questions across seven error categories. |
| Outcome: | The proposed datasets show that LLMs struggle with Korean commonsense reasoning . human accuracy benchmarked at approximately 85%, while GPT-4’s performance lags at about 74%, and other LLM models demonstrate an average accuracy of around 42%. |
KoLEG: On-the-Fly Korean Legal Knowledge Editing with Continuous Retrieval (2025.findings-emnlp)
Copied to clipboard
Jaehyung Seo, Dahyun Jung, Jaewook Lee, Yongchan Chun, Dongjun Kim, Hwijung Ryu, Donghoon Shin, Heuiseok Lim
| Challenge: | a recent study shows that Korean legal knowledge is subject to frequent temporal updates driven by societal needs and government policies. |
| Approach: | They propose a Korean Legal knowledge editing framework enhanced with continuous retrieval . they employ an Editing-Aware Learning Strategy and a LawEdit Retriever . |
| Outcome: | a new framework outperforms existing methods for updating legal knowledge in Korean . it maintains robust performance in sequential editing and is qualitatively validated by legal experts. |
KEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processing (2023.emnlp-main)
Copied to clipboard
| Challenge: | Conventional evaluation metrics for automatic speech recognition systems produce a singular aggregate score, which is insufficient for understanding specific system vulnerabilities. |
| Approach: | They propose to introduce the Korean Error Explainable Benchmark Dataset for ASR and Post-processing (KEBAP) this method enables a more balanced assessment encompassing speech recognition accuracy and user readability. |
| Outcome: | The proposed method enables a more balanced assessment encompassing speech recognition accuracy and user readability. |
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation (2026.acl-long)
Copied to clipboard
Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen
| Challenge: | Existing models lack cultural alignment across modalities and languages . a new framework to assess cultural awareness across linguistics and languages is needed . |
| Approach: | They propose a framework that integrates tri-modally aligned cultural benchmarks and a five-dimensional evaluation protocol to assess cross-country awareness disparities. |
| Outcome: | The proposed framework assesses cultural awareness disparities across modalities and languages . it is the first dataset aligned at the input level across text, image, and speech . |
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand (2026.acl-long)
Copied to clipboard
| Challenge: | Existing summarization systems struggle to address diverse linguistic and cognitive barriers among general readers. |
| Approach: | They propose a multi-agent framework that integrates template-based planning with an iterative feedback loop guided by simulated readers and domain expert revision to address comprehension barriers such as unknown terms, missing contexts, and confusing sentences. |
| Outcome: | The proposed framework improves readability and factuality across multiple datasets and human evaluations show that it is more accessible to a wide range of readers. |
CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme Ingredients (2023.emnlp-main)
Copied to clipboard
| Challenge: | Korean morphological variations present unique opportunities and challenges in natural language processing (NLP), necessitating an advanced understanding of morpheme-based sentence construction. |
| Approach: | They propose a method to replicate morphological transformations inherent in Korean sentences based on lexical and functional morphemes through generative data augmentation. |
| Outcome: | The proposed method improves performance in Korean multiple classification datasets without incurring external data usage. |
Priming Ancient Korean Neural Machine Translation (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have focused on the restoration and translation of historical languages. |
| Approach: | They propose to use two different stimuli to priming ancient-Korean NMT . they confirm the possibility of developing a human-centric model based on cognitive science . |
| Outcome: | The proposed model can be used to translate historical Korean documents using neural machine translation. |
Hyper-BTS Dataset: Scalability and Enhanced Analysis of Back TranScription (BTS) for ASR Post-Processing (2024.findings-eacl)
Copied to clipboard
Chanjun Park, Jaehyung Seo, Seolhwa Lee, Junyoung Son, Hyeonseok Moon, Sugyeong Eo, Chanhee Lee, Heuiseok Lim
| Challenge: | Automatic Speech Recognition (ASR) post-processing requires substantial amounts of data, requiring expensive phonetic transcription experts. |
| Approach: | They propose a "Hyper-BTS" dataset that is five times larger than prior studies . they propose criteria for categorizing error types within ASR post-processing . |
| Outcome: | The proposed method can generate ASR inputs from clean text using a text-to-speech system. |
Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent frontier-level LLMs have saturated many previously difficult benchmarks, leaving little room for further differentiation. |
| Approach: | They propose a benchmark to evaluate whether LLMs can execute string-matching NLP metrics by strictly following step-by-step instructions. |
| Outcome: | The proposed benchmarks show that they can perform step-by-step execution, instruction adherence, numerical computation, and long-range consistency in handling intermediate results. |
Generative Interpretation: Toward Human-Like Evaluation for Educational Question-Answer Pair Generation (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing evaluation methods often fail to produce objective results and favor high similarity to the ground-truth question-answer pairs. |
| Approach: | They propose an alternative approach to evaluate question-answer generation using Generative Interpretation (GI) GI outperforms existing evaluation methods in terms of human alignment . |
| Outcome: | The proposed approach outperforms existing evaluation methods in human alignment and shows comparable performance with GPT3.5, only with BART-large. |
Empirical Analysis of Noising Scheme based Synthetic Data Generation for Automatic Post-editing (2022.lrec-1)
Copied to clipboard
| Challenge: | Automatic post-editing (APE) is a research field that aims to correct errors in translated sentences regardless of the utilized machine translation system. |
| Approach: | They propose a method for automatically generating APE data based on a noising scheme from a parallel corpus. |
| Outcome: | The proposed method shows that depending on the type of noise, the noising scheme-based APE data generation may lead to inferior performance. |
Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | LLMs are prone to generate responses to instruction-formatted statements in an instinctive manner, rather than comprehending the underlying user intention within the given instructions. |
| Approach: | They propose to use an instruction-following capability benchmark to evaluate LLMs' instruction understanding capability. |
| Outcome: | The proposed benchmark analyzes the instruction understanding capability of large language models with four instruction candidates and a single candidate. |
PicTalky: Augmentative and Alternative Communication for Language Developmental Disabilities (2022.aacl-demo)
Copied to clipboard
| Challenge: | Existing software packages are expensive and difficult to use, and only provide simple functions. |
| Approach: | They propose an AI-based AAC system called PicTalky that can improve communication skills for children with language disabilities. |
| Outcome: | The proposed system improves communication skills and language comprehension abilities for children with language disabilities. |
Leveraging Pre-existing Resources for Data-Efficient Counter-Narrative Generation in Korean (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets and methods for detecting hate speech are limited by resource-intensive nature and only focus on the primary language. |
| Approach: | They propose a Korean Hate Speech Counter Punch (KHSCP) method that generates fact-based responses to hate speech in the Korean language and propose to use existing resources to overcome data scarcity. |
| Outcome: | The proposed method can overcome data scarcity in low-resource environments by leveraging existing resources. |
HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to document-based Opendomain Question Answering (ODQA) use flat text chunks or page-level images to locate the correct document. |
| Approach: | They propose a hierarchical tree-based multimodal retrieval framework that elevates document hierarchy to a first-class retrieval signal. |
| Outcome: | The proposed framework outperforms page- and chunk-based baselines on ODQA benchmarks and improves retrieval recall by 12.9% and end-to-end QA performance by 6.8%. |
The Impact of Negated Text on Hallucination with Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies on hallucination in large language models (LLMs) have been actively progressing in natural language processing. |
| Approach: | They propose to examine whether LLMs can recognize contextual shifts caused by negation and still reliably distinguish hallucinations comparable to affirmative cases. |
| Outcome: | The proposed model can detect hallucinations comparable to affirmative cases, but it is difficult to detect them in negated text, the authors show . |
QUAK: A Synthetic Quality Estimation Dataset for Korean-English Neural Machine Translation (2022.coling-1)
Copied to clipboard
| Challenge: | despite its high utility, there are limitations concerning manual QE data creation. |
| Approach: | They propose to generate a Korean-English QE dataset that is fully automatic . they find that the algorithm is more accurate and faster than manual QE . |
| Outcome: | The proposed datasets show that they scale up to 1.58M and 6.58M, respectively, and show that the results are significantly better when compared to the previous datasets. |
Intelligent Predictive Maintenance RAG framework for Power Plants: Enhancing QA with StyleDFS and Domain Specific Instruction Tuning (2024.emnlp-industry)
Copied to clipboard
Seongtae Hong, Joong Shin, Jaehyung Seo, Taemin Lee, Jeongbae Park, Cho Young, Byeongho Choi, Heuiseok Lim
| Challenge: | Existing off-premise Question-Answering systems based on Large Language Models face data leakage and domain-specific tuning challenges. |
| Approach: | They propose an on-premise intelligent PMS framework based on a chunking method . they propose instruction tuning using relevant domain-specific data improves LLM performance . |
| Outcome: | The proposed framework improves performance even under limited data conditions. |
MultiDocFusion : Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing text chunking methods neglect complex and long industrial document structures, causing information loss and reduced answer quality. |
| Approach: | They propose a multimodal chunking pipeline that detects document regions and extracts text from them via OCR. |
| Outcome: | Extensive tests show that MultiDocFusion improves retrieval precision by 8–15% and ANLS QA scores by 2–3% compared to baselines. |