Constructing a Dependency Treebank for Second Language Learners of Korean (2024.lrec-main)

Copied to clipboard

Challenge: a manually annotated syntactic treebank is available for second language learners . the dataset includes 7,530 sentences (66,982 words; 129,333 morphemes)
Approach: They propose to manually annotate syntactic treebanks based on Universal Dependencies from Korean written data.
Outcome: The proposed dataset includes 7,530 sentences and 129,333 morphemes from Korean learners.

Similar Papers

Building Universal Dependency Treebanks in Korean (L18-1)

Copied to clipboard

Challenge: Several treebanks were introduced for Korean, all of which comprised annotation of morphemes and phrase structure trees, each following its own set of guidelines.
Approach: They propose to use Korean treebanks as dependency trees and to analyze their performance using morpheme-level annotations.
Outcome: The Korean portion of the Google UD Treebank, the Penn Korean Treebank and the KAIST Treebank are re-tokenized and assessed for errors.
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)

Copied to clipboard

Challenge: Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages.
Approach: They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer .
Outcome: The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers.
Constructing Korean Learners’ L2 Speech Corpus of Seven Languages for Automatic Pronunciation Assessment (2024.lrec-main)

Copied to clipboard

Challenge: Multilingual L2 speech corpora for automatic speech assessment are currently available, but lack comprehensive annotations of L2 from non-native speakers of various languages.
Approach: They propose to use Korean learners’ L2 speech corpus of seven languages to develop automatic speech assessment.
Outcome: The proposed corpus contains 1,200 hours of L2 speech data from Korean learners (400 hours for English, 200 hours each for Japanese and Chinese, 100 hours each in French, German, Spanish, and Russian).
Korean L2 Vocabulary Prediction: Can a Large Annotated Corpus be Used to Train Better Models for Predicting Unknown Words? (L18-1)

Copied to clipboard

Challenge: a recent study suggests that a classifier trained on unknown words may yield better results for L2 learners.
Approach: They propose to use a supervised learning classifier to predict word complexity in Korean . they propose to train models on annotated corpus of unknown words with 71 % precision .
Outcome: The proposed model recalls 80 % of unknown words with 71 % precision.
Yet Another Format of Universal Dependencies for Korean (2022.coling-1)

Copied to clipboard

Challenge: Existing dependency parsers for Korean do not perform as well as their English counterparts due to the complexity of Korean's linguistic features.
Approach: They propose a morpheme-based Korean dependency parsing format and propose to adopt it to Universal Dependencies.
Outcome: The proposed format outperforms parsing results for Korean UD treebanks and detailed error analysis.
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)

Copied to clipboard

Challenge: spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all.
Approach: They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme.
Outcome: The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation.
A Universal Dependencies Treebank of Ancient Hebrew (2022.lrec-1)

Copied to clipboard

Challenge: Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew.
Approach: They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax.
Outcome: The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax.
The Persian Dependency Treebank Made Universal (2022.lrec-1)

Copied to clipboard

Challenge: Existing universal dependency treebanks are lacking sufficient annotated data.
Approach: They propose a method for converting Persian Dependency Treebank to Universal Dependencies using an automatic method.
Outcome: The proposed method is more compatible with Universal Dependencies than the Uppsala Persian Universal Dependency Treebank.
Semi-automatic Korean FrameNet Annotation over KAIST Treebank (L18-1)

Copied to clipboard

Challenge: Annotating FrameNet over raw sentences is an expensive and complex task, because of which we have designed a semi-automatic annotation approach.
Approach: They propose to use Korean FrameNet annotations to build a frame-semantic parser for English using full-text annotation and partially annotated exemplar sentences to train their models.
Outcome: The proposed model is based on a lexical database of the Korean FrameNet, and its current scope, status, and limitations are discussed in the paper.
L1-L2 Parallel Treebank of Learner Chinese: Overused and Underused Syntactic Structures (L18-1)

Copied to clipboard

Challenge: Currently, the treebank consists of 600 L2 sentences and 697 L1 sentences.
Approach: They propose to use "L1-L2 parallel treebanks" to facilitate analyses of learner language.
Outcome: The proposed treebank consists of 600 L2 sentences and 697 L1 sentences.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations