Papers by Kyubyong Park

5 papers
K-HATERS: A Hate Speech Detection Corpus in Korean with Target-Specific Ratings (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets on hate speech detection focus on overt forms of hate . however, a majority of these resources are English-centric, focusing on overtones of hate.
Approach: They propose a new corpus for hate speech detection in Korean with target-specific offensiveness ratings that offer a three-point Likert scale.
Outcome: The proposed corpus is the largest offensive language corpus in Korean and offers target-specific ratings on a three-point Likert scale.
Jejueo Datasets for Machine Translation and Speech Synthesis (2020.lrec-1)

Copied to clipboard

Challenge: Jejueo, or the Jeju language, is a minority language used on Jeju Island . there have been many efforts to revitalize the language, but few computational approaches have been used to solve its problems.
Approach: They construct two new Jejueo datasets using interviews and transcripts . they build machine translation and speech synthesis using these datasets based on their results .
Outcome: The proposed datasets will attract interest of both language and machine learning communities.
KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmark datasets for natural language inference and semantic textual similarity (STS) are not available in the Korean language.
Approach: They construct and release new datasets for Korean NLI and STS . they machine-translate existing English training sets and manually translate development and test sets into Korean to accelerate research on Korean NLU.
Outcome: The proposed datasets are available at https://github.com/kakaobrain/KorNLUDatasets.
word2word: A Collection of Bilingual Lexicons for 3,564 Language Pairs (2020.lrec-1)

Copied to clipboard

Challenge: Our dataset provides top-k word translations in 3,564 (directed) language pairs across 62 languages in OpenSubtitles2018.
Approach: They propose a dataset and an open-source Python package for cross-lingual word translations extracted from sentence-level parallel corpora.
Outcome: The proposed bilingual lexicons have high coverage and achieve competitive translation quality for several language pairs.
An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks (2020.aacl-main)

Copied to clipboard

Challenge: Traditionally, tokenization is the very first step in most text processing works.
Approach: They propose to use morphological segmentation followed by BPE for Korean NLP tasks . they empirically examine what is the best tokenization strategy for Korean to/from English .
Outcome: The proposed approach is best for Korean to/from English machine translation and natural language understanding tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations