Papers by Ohjoon Kwon

5 papers
SLM as Guardian: Pioneering AI Safety with Small Language Model (2024.emnlp-industry)

Copied to clipboard

Challenge: Prior safety research on large language models focused on aligning them to safety requirements, but internalizing such safeguard features into larger models brought challenges of higher training cost and unintended degradation of helpfulness.
Approach: They propose a multi-task learning mechanism that integrates harmful query detection and safeguard response into a single model.
Outcome: The proposed approach outperforms the publicly available LLMs in harmful query detection and safeguard response generation.
Taxonomy and Analysis of Sensitive User Queries in Generative AI Search System (2025.findings-naacl)

Copied to clipboard

Challenge: generative LLMs have been used by industries for various purposes, but limited resources and limited experience hinder their deployment and maintenance.
Approach: They propose a taxonomy for sensitive search queries and outline approaches to generating generative LLMs.
Outcome: The proposed model can be used to analyze sensitive queries from real users.
QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines (2025.acl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have been widely used for relevance assessment in information retrieval, but maintaining and updating such models is resource-intensive, limiting their feasibility in dynamic and multilingual search environments.
Approach: They propose to combine a generative SLM with an embedding-based SLM to achieve higher relevance judgment accuracy while reducing computational costs.
Outcome: The proposed approach outperforms state-of-the-art LLMs in relevance assessment tasks while reducing computational costs.
Conflict and Overlap Classification in Construction Standards Using a Large Language Model (2025.naacl-industry)

Copied to clipboard

Challenge: Current manual approaches to analyzing overlapping or conflicting content are time-consuming, costly, and error-prone.
Approach: They propose a large language model that uses a construction domain-adapted large language for the semantic comparison of sentences in construction standards.
Outcome: The proposed framework achieves 97.9% accuracy and 0.907 macro F1-score in classifying sentences from Korean construction standards as overlapping, conflicting, or neutral.
Handling Out-Of-Vocabulary Problem in Hangeul Word Embeddings (2021.eacl-main)

Copied to clipboard

Challenge: Word embedding is considered an essential factor in improving the performance of various Natural Language Processing (NLP) models.
Approach: They propose a Hangeul word embedding model that infers original word embeds from typos while maintaining high performance.
Outcome: The proposed model performs well against typos while maintaining high performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations