Papers by Donghyeon Lee
SLM as Guardian: Pioneering AI Safety with Small Language Model (2024.emnlp-industry)
Copied to clipboard
Ohjoon Kwon, Donghyeon Jeon, Nayoung Choi, Gyu-Hwung Cho, Hwiyeol Jo, Changbong Kim, Hyunwoo Lee, Inho Kang, Sun Kim, Taiwoo Park
| Challenge: | Prior safety research on large language models focused on aligning them to safety requirements, but internalizing such safeguard features into larger models brought challenges of higher training cost and unintended degradation of helpfulness. |
| Approach: | They propose a multi-task learning mechanism that integrates harmful query detection and safeguard response into a single model. |
| Outcome: | The proposed approach outperforms the publicly available LLMs in harmful query detection and safeguard response generation. |
PURE: Post-hoc Unlocking and REfinement for Discrete Diffusion Decoding (2026.findings-acl)
Copied to clipboard
| Challenge: | Masked diffusion language models (MDLMs) are limited by a monotonic unmasking policy, where committed tokens cannot be revised. |
| Approach: | They propose a training-free inference algorithm for two-phase decoding that unlocks unstable regions through deterministic window masking and stochastic leftward relaxation. |
| Outcome: | The proposed algorithm significantly improves accuracy on reasoning benchmarks on GSM8K. |
ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing evaluation methods do not assess whether large language models fully utilize contextual information. |
| Approach: | They introduce a new metric to assess LLMs' ability to fully utilize contextual information. |
| Outcome: | The proposed benchmark comprises 1,986 test instances spanning four long-context tasks with high IC scores in the domains of books, debates, medicine, and law. |
From Generation to Selection: Findings of Converting Analogical Problem-Solving into Multiple-Choice Questions (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Abstract and Reasoning Corpus (ARC) is a benchmark designed to evaluate reasoning abilities alone by reducing the amount of prior knowledge and data required to solve the tasks. |
| Approach: | They propose a multiple-choice format suitable for assessing stages like Understand and Apply in Large Language Models (LLMs). |
| Outcome: | The proposed model supports analogical reasoning and evidence analysis, but LLMs use shortcuts in the MC-LARC format. |
Taxonomy and Analysis of Sensitive User Queries in Generative AI Search System (2025.findings-naacl)
Copied to clipboard
Hwiyeol Jo, Taiwoo Park, Hyunwoo Lee, Nayoung Choi, Changbong Kim, Ohjoon Kwon, Donghyeon Jeon, Eui Hyeon Lee, Kyoungho Shin, Lim Sun Suk, Kyungmi Kim, Lee Jihye, Sun Kim
| Challenge: | generative LLMs have been used by industries for various purposes, but limited resources and limited experience hinder their deployment and maintenance. |
| Approach: | They propose a taxonomy for sensitive search queries and outline approaches to generating generative LLMs. |
| Outcome: | The proposed model can be used to analyze sensitive queries from real users. |
MAFiD: Moving Average Equipped Fusion-in-Decoder for Question Answering over Tabular and Textual Data (2023.findings-eacl)
Copied to clipboard
| Challenge: | Experimental results show that Transformer-based questions have a "long" hybrid sequence over tabular and textual elements, causing long-range reasoning problems. |
| Approach: | They propose a moving average-equipped fusion-in-decoder to handle long-range reasoning problems . they use FiD and EMA to combine different levels of reasoning . |
| Outcome: | Experimental results show that the proposed model increases exact matching and F1 by 1.1 and 1.7 on the blind test set. |
From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines (2026.acl-industry)
Copied to clipboard
| Challenge: | Existing methods that optimize for relevance overlook document trustworthiness . Generative information retrieval (GenIR) is a promising paradigm for retrieval tasks . |
| Approach: | They propose an Authority-aware Generative Retriever (AuthGR) that incorporates authority into GenIR. |
| Outcome: | The proposed framework improves authority and accuracy in real-world user engagement and reliability. |
PRIME: Ultra-Low-Rank Principal–Residual Model Merging (2026.findings-acl)
Copied to clipboard
Seung-Ho Lee, Kyungsu Lee, Bazarvaani Zuchi, Jeongmin Ahn, Insuk Seo, Donghyeon Jeon, Inho Kang, Seung-Hoon Na
| Challenge: | Existing methods for model merging have been limited by task-specific performance and task-related tasks. |
| Approach: | They propose an ultra-low-rank principal-residual model merging framework that decomposes task vector merging into two stages. |
| Outcome: | Experiments on eight natural language processing tasks show that PRIME outperforms existing models while preserving the task-specific capabilities of the original models. |
QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines (2025.acl-industry)
Copied to clipboard
| Challenge: | Large language models (LLMs) have been widely used for relevance assessment in information retrieval, but maintaining and updating such models is resource-intensive, limiting their feasibility in dynamic and multilingual search environments. |
| Approach: | They propose to combine a generative SLM with an embedding-based SLM to achieve higher relevance judgment accuracy while reducing computational costs. |
| Outcome: | The proposed approach outperforms state-of-the-art LLMs in relevance assessment tasks while reducing computational costs. |
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing function-calling benchmarks focus on single-turn interactions but ignore complexity of real-world scenarios. |
| Approach: | They propose a framework that constructs practical function-calling datasets by synthesizing conversations through a tool graph that maintains dependencies across rounds. |
| Outcome: | The proposed framework synthesizes conversations through a tool graph that maintains dependencies across rounds and a multi-agent system with distinct personas to enhance dialogue naturalness. |