Papers by Younggyun Hahm
Automatic Wordnet Mapping: from CoreNet to Princeton WordNet (L18-1)
Copied to clipboard
| Challenge: | Existing mappings focus on identifying the semantic categories of CoreNet, but not the word senses. |
| Approach: | They propose to map the word senses of CoreNet into Princeton WordNet synsets by lexical relations by a taxonomy. |
| Outcome: | The proposed mapping bridging the gap between CoreNet and WordNet shows that the word senses of CoreNet are mapped with precision of 91.2%. |
Semi-automatic Korean FrameNet Annotation over KAIST Treebank (L18-1)
Copied to clipboard
| Challenge: | Annotating FrameNet over raw sentences is an expensive and complex task, because of which we have designed a semi-automatic annotation approach. |
| Approach: | They propose to use Korean FrameNet annotations to build a frame-semantic parser for English using full-text annotation and partially annotated exemplar sentences to train their models. |
| Outcome: | The proposed model is based on a lexical database of the Korean FrameNet, and its current scope, status, and limitations are discussed in the paper. |
Unsupervised Korean Word Sense Disambiguation using CoreNet (L18-1)
Copied to clipboard
| Challenge: | Unsupervised learning based Korean word sense disambiguation is needed to distinguish between sense candidates. |
| Approach: | They investigated unsupervised Korean word sense disambiguation using CoreNet, a Korean lexical semantic network. |
| Outcome: | The proposed method exhibited an 80.9% accuracy on the datasets constructed and proved to be effective for practical applications. |
Crowdsourcing in the Development of a Multilingual FrameNet: A Case Study of Korean FrameNet (2020.lrec-1)
Copied to clipboard
| Challenge: | Using current methods, the construction of multilingual FrameNets is expensive and complex. |
| Approach: | They evaluated whether crowdsourcing approaches captured cross-cultural and cross-linguistic meanings . they found that crowd workers made intuitive choices comparable to trained FrameNet experts . |
| Outcome: | The results are now available in Korean FrameNet 1.1. |
Optimizing Language Augmentation for Multilingual Large Language Models: A Case Study on Korean (2024.lrec-main)
Copied to clipboard
ChangSu Choi, Yongbin Jeong, Seoyoon Park, Inho Won, HyeonSeok Lim, SangMin Kim, Yejee Kang, Chanhyuk Yoon, Jaewan Park, Yiseul Lee, HyeJin Lee, Younggyun Hahm, Hansaem Kim, KyungTae Lim
| Challenge: | Large language models (LLMs) use pretraining to predict the subsequent word, but less-resourced languages are being overlooked. |
| Approach: | They propose to expand the MLLM vocabularies to enhance expressiveness and use bilingual data for pretraining to align the high- and less-resourced languages. |
| Outcome: | The proposed model outperforms existing models in qualitative analyses compared to Korean monolingual models. |