Papers by Suwan Kim

2 papers
ToDi: Token-wise Distillation via Fine-Grained Divergence Control (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) offer impressive performance but are impractical for resource-constrained deployment due to high latency and energy consumption.
Approach: They propose a method that adaptively combines FKL and RKL per token using a sigmoid-based weighting function derived from the teacher-student probability log-ratio.
Outcome: The proposed method outperforms baselines using uniform or less granular strategies across instruction-following benchmarks.
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing evaluation tools rely on translations of English datasets or translation-specific benchmarks such as WMT 21 to assess large language models.
Approach: They propose a dataset curated to challenge models lacking Korean cultural and contextual depth.
Outcome: The HAE-RAE Bench challenges models lacking Korean cultural and contextual depth by highlighting their aptitude for recalling Korean-specific knowledge and cultural contexts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations