Papers by Minseon Kim
Language Detoxification with Attribute-Discriminative Latent Space (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods to detoxify toxic text require excessive memory, computations and time. |
| Approach: | They propose a method to generate toxic text using an attribute-discriminative latent space. |
| Outcome: | The proposed method outperforms baselines on detoxified language and dialogue generation tasks while being time- and memory-efficient. |
FLUID QA: A Multilingual Benchmark for Figurative Language Usage in Dialogue across English, Chinese, and Korean (2025.emnlp-main)
Copied to clipboard
| Challenge: | Figurative language is a core component of everyday communication . existing benchmarks focus on sentence-level classification or inference tasks . |
| Approach: | They propose a multilingual benchmark that evaluates figurative usage in dialogue . they use a sentence-level diagnostic task to embed figurativ choices into multi-turn contexts . |
| Outcome: | The benchmark evaluates large language models' ability to use figurative expressions coherently in conversation. |
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings (2026.eacl-industry)
Copied to clipboard
Jean-Philippe Corbeil, Minseon Kim, Maxime Griot, Sheela Agarwal, Alessandro Sordoni, Francois Beaulieu, Paul Vozila
| Challenge: | Existing risk evaluations focused on general safety benchmarks, resulting in role-dependent vulnerabilities in real-world medical and clinical deployments. |
| Approach: | They propose a patient-oriented dataset called PatientSafetyBench that evaluates a variety of open- and closed-source LLMs. |
| Outcome: | The proposed benchmark examines medical risks from 466 open- and closed-source LLMs across 5 risk categories. |