Towards Standardized Annotation and Parsing for Korean FrameNet (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on Korean FrameNet have focused on English, but annotations are not optimally designed for Korean.
Approach: They propose a morphologically enhanced annotation strategy for Korean FrameNet datasets and parsing by leveraging the CoNLL-U format.
Outcome: The proposed method improves the annotation accuracy of Korean FrameNet datasets and their parsers.

Similar Papers

Semi-automatic Korean FrameNet Annotation over KAIST Treebank (L18-1)

Copied to clipboard

Challenge: Annotating FrameNet over raw sentences is an expensive and complex task, because of which we have designed a semi-automatic annotation approach.
Approach: They propose to use Korean FrameNet annotations to build a frame-semantic parser for English using full-text annotation and partially annotated exemplar sentences to train their models.
Outcome: The proposed model is based on a lexical database of the Korean FrameNet, and its current scope, status, and limitations are discussed in the paper.
Yet Another Format of Universal Dependencies for Korean (2022.coling-1)

Copied to clipboard

Challenge: Existing dependency parsers for Korean do not perform as well as their English counterparts due to the complexity of Korean's linguistic features.
Approach: They propose a morpheme-based Korean dependency parsing format and propose to adopt it to Universal Dependencies.
Outcome: The proposed format outperforms parsing results for Korean UD treebanks and detailed error analysis.
K-UniMorph: Korean Universal Morphology and its Feature Schema (2023.findings-acl)

Copied to clipboard

Challenge: Previously, the Korean language has been underrepresented in the field of morphological paradigms amongst hundreds of diverse world languages.
Approach: They propose a new Universal Morphology dataset for Korean that preserves its distinct characteristics.
Outcome: The proposed dataset extracts inflected Korean verb forms from the largest annotated corpus for Korean.
OpenKorPOS: Democratizing Korean Tokenization with Voting-Based Open Corpus Annotation (2022.lrec-1)

Copied to clipboard

Challenge: Korean uses spaces at larger-than-word boundaries, unlike other East-Asian languages.
Approach: They propose to use Korean morphological analyzers to provide a sequence of morpheme-level tokens, losing information in the tokenization process.
Outcome: The proposed scheme improves existing tagging scheme and makes it friendlier to generative tasks.
A Linguistically-Informed Annotation Strategy for Korean Semantic Role Labeling (2024.lrec-main)

Copied to clipboard

Challenge: Semantic role labeling is an essential component of semantic and syntactic processing of natural languages.
Approach: They propose an annotation strategy for Korean semantic role labeling that is in line with the previously proposed linguistic theories as well as the distinct properties of the Korean language.
Outcome: The proposed annotation strategy is consistent with the proposed linguistic theories and the distinct properties of the Korean language.
CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing (L18-1)

Copied to clipboard

Challenge: Using the universal dependencies framework, we address the need for a universal representation of morphological analysis that can capture alternative morphology of surface tokens and is compatible with the segmentation and morphologic annotation guidelines prescribed for UD treebanks.
Approach: They propose a new annotation format for word lattices that represent morphological analyses and a resource that obeys this format for a range of typologically different languages.
Outcome: The proposed model can capture alternative morphological analyses of surface tokens and is compatible with the segmentation and morphology guidelines prescribed for UD treebanks.
Transfer of Frames from English FrameNet to Construct Chinese FrameNet: A Bilingual Corpus-Based Approach (L18-1)

Copied to clipboard

Challenge: Current publicly available Chinese FrameNet has a relatively low coverage of frames and lexical units compared with other languages.
Approach: They propose an automatic way to construct Chinese FrameNet using a sentence-aligned English-Chinese bilingual corpus.
Outcome: The proposed resource can provide frame recommendations acceptable by annotators.
Rich Character-Level Information for Korean Morphological Analysis and Part-of-Speech Tagging (C18-1)

Copied to clipboard

Challenge: Korean is a highly agglutinative, character-rich language, requiring dictionary-less morphological analysis . a novel model can perform morphology and part-of-speech tagging without prior knowledge .
Approach: They propose a multi-stage action-based model that performs morphological transformation and part-of-speech tagging using arbitrary units of input.
Outcome: The proposed model achieves state-of-the-art word and sentence-level tagging accuracy with Korean corpus.
Constructing a Dependency Treebank for Second Language Learners of Korean (2024.lrec-main)

Copied to clipboard

Challenge: a manually annotated syntactic treebank is available for second language learners . the dataset includes 7,530 sentences (66,982 words; 129,333 morphemes)
Approach: They propose to manually annotate syntactic treebanks based on Universal Dependencies from Korean written data.
Outcome: The proposed dataset includes 7,530 sentences and 129,333 morphemes from Korean learners.
Building Universal Dependency Treebanks in Korean (L18-1)

Copied to clipboard

Challenge: Several treebanks were introduced for Korean, all of which comprised annotation of morphemes and phrase structure trees, each following its own set of guidelines.
Approach: They propose to use Korean treebanks as dependency trees and to analyze their performance using morpheme-level annotations.
Outcome: The Korean portion of the Google UD Treebank, the Penn Korean Treebank and the KAIST Treebank are re-tokenized and assessed for errors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations