Papers by Jaewook Lee
StepKE: Stepwise Knowledge Editing for Multi-Hop Question Answering (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing knowledge editing methods overlook interplay with pre-existing knowledge, leading to inconsistent edit propagation. |
| Approach: | stepKE integrates edited and existing knowledge for coherent multi-hop reasoning . stepKE decomposes multi-step questions into sequential single-hop sub-questions . |
| Outcome: | Experiments show that StepKE generates more accurate and consistent responses than baselines. |
Do Large Language Models Have “Emotion Neurons”? Investigating the Existence and Role (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations of LLMs' emotional capabilities have been criticized for not illuminating how emotion information is processed and represented within an LLM. |
| Approach: | They examine whether there are “emotion neurons” within large language models that selectively process and express certain emotions and what functional role they play. |
| Outcome: | The proposed model is based on the representative emotion theory of the six basic emotions and demonstrates that it is functionally significant to examine whether the prediction accuracy for a specific emotion decreases when the neurons are removed. |
Exploring Automated Keyword Mnemonics Generation with Large Language Models via Overgenerate-and-Rank (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Typically, creating verbal cues requires extensive human effort and is quite time-consuming. |
| Approach: | They propose a method for overgenerating and ranking verbal cues by prompting large language models to generate them and ranking them according to psycholinguistic measures and takeaways from a pilot user study. |
| Outcome: | The proposed method is comparable to human-generated mnemonics in imageability, coherence, and perceived usefulness, but there remains room for improvement due to the diversity in background and preference among language learners. |
Small Changes, Big Impact: How Manipulating a Few Neurons Can Drastically Alter LLM Aggression (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models have led to innovations in various domains such as education, healthcare, and finance, while raising serious concerns that they can be easily misused for malicious purposes. |
| Approach: | They identify specific neurons (“aggression neurons”) closely related to the expression of aggression and analyze how manipulating them affects the model’s overall aggression. |
| Outcome: | The proposed model outputs show that manipulating neurons can increase aggression by up to 33% in all models and even more extreme when they are concentrated in certain layers. |
CoME: An Unlearning-based Approach to Conflict-free Model Editing (2025.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) often retain outdated or incorrect information from pre-training, which undermines their reliability. |
| Approach: | They propose a conflict-free model editing framework that selectively removes outdated knowledge from LLMs to improve their accuracy and reliability. |
| Outcome: | The proposed framework improves both editing accuracy and model reliability when applied to existing editing methods. |
KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Language models are striving to grasp commonsense reasoning, but they are lacking in Korean commons- ense benchmarks. |
| Approach: | They present a fine-grained benchmark dataset focused on Korean commonsense reasoning that includes multiple-choice questions across seven error categories. |
| Outcome: | The proposed datasets show that LLMs struggle with Korean commonsense reasoning . human accuracy benchmarked at approximately 85%, while GPT-4’s performance lags at about 74%, and other LLM models demonstrate an average accuracy of around 42%. |
KoLEG: On-the-Fly Korean Legal Knowledge Editing with Continuous Retrieval (2025.findings-emnlp)
Copied to clipboard
Jaehyung Seo, Dahyun Jung, Jaewook Lee, Yongchan Chun, Dongjun Kim, Hwijung Ryu, Donghoon Shin, Heuiseok Lim
| Challenge: | a recent study shows that Korean legal knowledge is subject to frequent temporal updates driven by societal needs and government policies. |
| Approach: | They propose a Korean Legal knowledge editing framework enhanced with continuous retrieval . they employ an Editing-Aware Learning Strategy and a LawEdit Retriever . |
| Outcome: | a new framework outperforms existing methods for updating legal knowledge in Korean . it maintains robust performance in sequential editing and is qualitatively validated by legal experts. |
PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Vocabulary acquisition is a challenge for second-language learners when learning typologically distant languages such as English and Korean, where phonological and structural mismatches complicate vocabulary learning. |
| Approach: | They propose a cross-lingual mnemonic generation system that performs IPA-based phonological adaptation and syllable-aware alignment to retrieve L1 keyword sequence and uses LLMs to generate verbal cues. |
| Outcome: | The proposed system outperforms human-written and automated mnemonics in a short-term recall test with human participants and achieves quality comparable to human-writing mnms. |
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation protocols for text generation suffer from rating inconsistencies . lexical overlap-based metrics align poorly with human judgments . |
| Approach: | They propose a checklist-based evaluation framework that improves rating reliability via decomposed binary questions. |
| Outcome: | The proposed framework improves rating reliability by decomposing binary questions . it improves agreement across evaluator models by 0.45 and reduces score variance . human evaluation remains the gold standard, but it #, Equal contribution. |
What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers (2021.emnlp-main)
Copied to clipboard
Boseop Kim, HyoungSeok Kim, Sang-Woo Lee, Gichang Lee, Donghyun Kwak, Jeon Dong Hyeon, Sunghyun Park, Sungju Kim, Seonhoon Kim, Dongpil Seo, Heungsub Lee, Minyoung Jeong, Sungjae Lee, Minsub Kim, Suk Hyun Ko, Seokhun Kim, Taeyong Park, Jinuk Kim, Soyoung Kang, Na-Hyeon Ryu, Kang Min Yoo, Minsuk Chang, Soobin Suh, Sookyo In, Jinseong Park, Kyungduk Kim, Hiun Kim, Jisu Jeong, Yong Goo Yeo, Donghoon Ham, Dongju Park, Min Young Lee, Jaewook Kang, Inho Kang, Jung-Woo Ha, Woomyoung Park, Nako Sung
| Challenge: | GPT-3 has been used to train large-scale language models on hundreds of billion scale data. |
| Approach: | They propose a Korean variant of GPT-3 that uses Korean tokens to train in-context models. |
| Outcome: | The proposed method shows state-of-the-art zero-shot and few-shot learning on downstream tasks in Korean. |
Does the Emotional Understanding of LVLMs Vary Under High-Stress Environments and Across Different Demographic Attributes? (2025.acl-long)
Copied to clipboard
| Challenge: | According to psychological and neuroscientific research, a high-stress environment can restrict attentional resources and intensify negative affect, thereby impairing the ability to understand emotions. |
| Approach: | They constructed a large-vision language model that combines race, gender, and age group and used the Pretend prompt technique to induce LVLMs to interpret others’ emotions. |
| Outcome: | The results suggest that the effects of high-stress and demographic attributes identified in human research may also be reflected in LVLMs. |
CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme Ingredients (2023.emnlp-main)
Copied to clipboard
| Challenge: | Korean morphological variations present unique opportunities and challenges in natural language processing (NLP), necessitating an advanced understanding of morpheme-based sentence construction. |
| Approach: | They propose a method to replicate morphological transformations inherent in Korean sentences based on lexical and functional morphemes through generative data augmentation. |
| Outcome: | The proposed method improves performance in Korean multiple classification datasets without incurring external data usage. |
Enhancing Out-of-Distribution Detection in Natural Language Understanding via Implicit Layer Ensemble (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Out-of-distribution (OOD) detection aims to discern outliers from the intended data distribution, which is crucial to maintaining high reliability and a good user experience. |
| Approach: | They propose a framework that encourages intermediate features to learn layer-specialized representations and assembles them implicitly into a single representation to absorb rich information in the pre-trained language model. |
| Outcome: | The proposed framework is significantly more effective than previous studies in intent classification and OOD datasets. |
Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximization (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for mnemonic generation in Japanese are limited in their interpretability due to script differences. |
| Approach: | They propose a method that models the mnemonic construction process as driven by common rules. |
| Outcome: | The proposed method performs well in the cold-start setting for new learners while providing insight into the mechanisms behind effective mnemonic creation. |
Generative Interpretation: Toward Human-Like Evaluation for Educational Question-Answer Pair Generation (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing evaluation methods often fail to produce objective results and favor high similarity to the ground-truth question-answer pairs. |
| Approach: | They propose an alternative approach to evaluate question-answer generation using Generative Interpretation (GI) GI outperforms existing evaluation methods in terms of human alignment . |
| Outcome: | The proposed approach outperforms existing evaluation methods in human alignment and shows comparable performance with GPT3.5, only with BART-large. |
Make LLMs See Like Investigators, Not Just Think More: The Role of Structured Analysis in Investigative Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Criminal investigators and intelligence analysts have developed structured analytic techniques to evaluate competing hypotheses under incomplete information. |
| Approach: | They focus on the task of analyzing evidence from complex narratives and identifying the perpetrator among suspects using the MuSR murder mystery benchmark. |
| Outcome: | The PRISM framework outperforms general-purpose strategies across all models, with its effectiveness manifesting regardless of model scale. |
Towards Scalable Lifelong Knowledge Editing with Selective Knowledge Suppression (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods to modify knowledge are limited due to high training costs and lack stability during sequential edits due to catastrophic forgetting. |
| Approach: | They propose a framework to modify specific knowledge of large language models without retraining the entire model. |
| Outcome: | Extensive experiments on ZSRE, Counterfact, and RIPE show that LightEdit outperforms existing lifelong knowledge editing methods. |
Analyzing Key Factors Influencing Emotion Prediction Performance of VLLMs in Conversational Contexts (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that large language models and vision large language model (VLLMs) possess EI and the ability to understand emotional stimuli in the form of text and images. |
| Approach: | They analyze the key elements affecting the emotion prediction performance of VLLMs in conversational contexts. |
| Outcome: | The proposed model performance was compared with other models in a conversational context. |
Exploring Automated Distractor Generation for Math Multiple-choice Questions via Large Language Models (2024.findings-naacl)
Copied to clipboard
Wanyong Feng, Jaewook Lee, Hunter McNichols, Alexander Scarlatos, Digory Smith, Simon Woodhead, Nancy Ornelas, Andrew Lan
| Challenge: | Multiple-choice questions (MCQs) are easy to administer and grade . but crafting high-quality distractors remains labor-intensive and limited scalability . |
| Approach: | They propose to automate the generation of distractors in math MCQs by using large language models to generate distractors. |
| Outcome: | The proposed methods can generate valid distractors, but they are less adept at anticipating common errors or misconceptions among real students. |
Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models (2025.naacl-industry)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have impacted the writing process, enhancing productivity by collaborating with humans in content creation platforms. |
| Approach: | They propose a framework that uses explicit outlines to guide LLMs in generating goal-oriented, high-quality text. |
| Outcome: | The proposed approach significantly improves text quality according to evaluations by LLMs and professional writers. |
GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies report that prompt-based direct classification eliminates the need for fine-tuning but lacks data and inference scalability. |
| Approach: | They propose a data augmentation technique that leverages large-scale language models to generate real text samples from a mixture of real samples. |
| Outcome: | The proposed method outperforms existing methods on diverse classification tasks. |
HyperT5: Towards Compute-Efficient Korean Language Modeling (2023.acl-industry)
Copied to clipboard
| Challenge: | Pretraining and fine-tuning language models is a common practice in NLP, but deploying general-purpose language models without the abundant computation or data resources is proving difficult. |
| Approach: | They propose a sequence-to-sequence language model architecture that can be more practical and compute-efficient than the decoder-oriented approach. |
| Outcome: | The proposed language model outperforms competing models in Korean benchmarks and is more efficient in low-resource settings. |
Can We Entrust Justice to AI?: How Persona Traps Contaminate Reasoning in Criminal Investigation (2026.findings-acl)
Copied to clipboard
| Challenge: | a study injected personas into neutralized murder mystery scenarios to examine reasoning stability of large language models . adrian s. gupta: if models are used to evaluate suspects, are they free from the trap of implicit bias? he says current alignment techniques focus on identity-based bias while neglecting relationship-based ones . |
| Approach: | a study systematically injected personas into neutralized murder mystery scenarios . it found that implicit bias propagation was observed across all models . authors propose stability evaluation should encompass outputs and reasoning processes . |
| Outcome: | The proposed pipeline can analyze evidence and evaluate suspects in murder mysteries . it shows that models outwardly state "that information is irrelevant to the judgment" the proposed pipeline could be extended to include reasoning processes, authors say . |
Safeguarding Privacy of Retrieval Data against Membership Inference Attacks: Is This Query Too Close to Home? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | delivering private retrieved documents directly to LLMs introduces vulnerability to membership inference attacks . |
| Approach: | They propose a similarity-based membership inference attack detection framework for RAG . they propose obfuscate attackers, maintain data utility, and remain system-agnostic . |
| Outcome: | The proposed framework can detect and hide membership inference attacks, while remaining system-agnostic against them. |
Simulated Students in Tutoring Dialogues: Substance or Illusion? (2026.acl-long)
Copied to clipboard
| Challenge: | evaluating the effectiveness of new technology requires real students, which is time-consuming and hard to scale up. |
| Approach: | They propose to define the student simulation task and benchmark a wide range of student simulation methods on these metrics. |
| Outcome: | The proposed evaluation metrics show that prompting strategies perform poorly on a real-world tutoring dialogue dataset. |
A Framework for Vision-Language Warm-up Tasks in Multimodal Dialogue Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for building multimodal open-domain dialogue agents based on large datasets are limited in real-world settings . |
| Approach: | They propose a new learning strategy called vision-language warm-up tasks for multimodal dialogue models that relies solely on learning from target data. |
| Outcome: | The proposed learning strategy achieves comparable and in some cases superior performance compared to existing state-of-the-art models on various evaluation metrics. |