Papers by Seungho Lee
ToxiPrompt: A Two-Stage Red-Teaming Approach for Balancing Adversarial Prompt Diversity and Response Toxicity (2026.eacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) pose safety risks, but current redteaming methods rely on human testers manually designing adversarial prompts. |
| Approach: | They propose a red-teaming method that generates adversarial prompts to elicit unsafe behavior of target LLMs. |
| Outcome: | The proposed approach outperforms state-of-the-art methods in diversity and toxicity . it performs well for multiple instruction-tuned target LLMs without re-tuning . |
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge (2024.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) must possess an understanding of the nation’s culture and basic knowledge. |
| Approach: | They propose to construct a national alignment benchmark, KorNAT, which measures the alignment between an LLM and a targeted country from two perspectives: social value alignment and common knowledge alignment. |
| Outcome: | The proposed model passes the national alignment score of 7 LLMs, indicating there is room for improvement. |