Papers by Younggyun Hahm

5 papers
Automatic Wordnet Mapping: from CoreNet to Princeton WordNet (L18-1)

Copied to clipboard

Challenge: Existing mappings focus on identifying the semantic categories of CoreNet, but not the word senses.
Approach: They propose to map the word senses of CoreNet into Princeton WordNet synsets by lexical relations by a taxonomy.
Outcome: The proposed mapping bridging the gap between CoreNet and WordNet shows that the word senses of CoreNet are mapped with precision of 91.2%.
Semi-automatic Korean FrameNet Annotation over KAIST Treebank (L18-1)

Copied to clipboard

Challenge: Annotating FrameNet over raw sentences is an expensive and complex task, because of which we have designed a semi-automatic annotation approach.
Approach: They propose to use Korean FrameNet annotations to build a frame-semantic parser for English using full-text annotation and partially annotated exemplar sentences to train their models.
Outcome: The proposed model is based on a lexical database of the Korean FrameNet, and its current scope, status, and limitations are discussed in the paper.
Unsupervised Korean Word Sense Disambiguation using CoreNet (L18-1)

Copied to clipboard

Challenge: Unsupervised learning based Korean word sense disambiguation is needed to distinguish between sense candidates.
Approach: They investigated unsupervised Korean word sense disambiguation using CoreNet, a Korean lexical semantic network.
Outcome: The proposed method exhibited an 80.9% accuracy on the datasets constructed and proved to be effective for practical applications.
Crowdsourcing in the Development of a Multilingual FrameNet: A Case Study of Korean FrameNet (2020.lrec-1)

Copied to clipboard

Challenge: Using current methods, the construction of multilingual FrameNets is expensive and complex.
Approach: They evaluated whether crowdsourcing approaches captured cross-cultural and cross-linguistic meanings . they found that crowd workers made intuitive choices comparable to trained FrameNet experts .
Outcome: The results are now available in Korean FrameNet 1.1.
Optimizing Language Augmentation for Multilingual Large Language Models: A Case Study on Korean (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) use pretraining to predict the subsequent word, but less-resourced languages are being overlooked.
Approach: They propose to expand the MLLM vocabularies to enhance expressiveness and use bilingual data for pretraining to align the high- and less-resourced languages.
Outcome: The proposed model outperforms existing models in qualitative analyses compared to Korean monolingual models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations