| Challenge: | a pilot study of Russian learner data with syntactic dependency relations is presented . a focus of recent work in the NLP community has been on grammar errors . |
| Approach: | They propose to annotate Russian learner data with syntactic dependency relations using a subset of sentences from two error-corrected Russian learners. |
| Outcome: | The proposed annotations are performed on a subset of Russian learner datasets. |
Similar Papers
Russian Learner Corpus: Towards Error-Cause Annotation for L2 Russian (2024.lrec-main)
Copied to clipboard
Daniil Kosakin, Sergei Obiedkov, Ivan Smirnov, Ekaterina Rakhilina, Anastasia Vyrenkova, Ekaterina Zalivina
| Challenge: | Russian Learner Corpus (RLC) is a large collection of learner texts written by native speakers of over forty languages. |
| Approach: | They propose an automatic error annotation tool that locates and labels errors according to a simplified version of the RLC error-type system. |
| Outcome: | The proposed tool locates and labels errors according to a simplified version of the RLC error-type system. |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
New Dataset and Strong Baselines for the Grammatical Error Correction of Russian (2021.findings-acl)
Copied to clipboard
| Challenge: | a new resource is created to evaluate grammatical error correction models in English . a subset of the dataset is annotated in Russian, which is hard to come by and expensive to annotate . |
| Approach: | They develop an annotated learner corpus of Russian extracted from the Lang-8 website. |
| Outcome: | The proposed dataset is compared against two state-of-the-art grammatical error correction models . the results show that the created corpus is more diverse than the existing one . |
Towards the Conversion of National Corpus of Polish to Universal Dependencies (2020.lrec-1)
Copied to clipboard
| Challenge: | a paper aims at enriching the manually annotated part of National Corpus of Polish with a syntactic layer. |
| Approach: | They enrich manually annotated part of Polish National Corpus with a syntactic layer and a UD dependency graph. |
| Outcome: | The proposed model outperforms a model trained on a smaller set of gold-standard trees in predicting part-of-speech tags, morphological features, lemmata and labelled dependency trees. |
Universal Dependencies Version 2 for Japanese (L18-1)
Copied to clipboard
Masayuki Asahara, Hiroshi Kanayama, Takaaki Tanaka, Yusuke Miyao, Sumire Uematsu, Shinsuke Mori, Yuji Matsumoto, Mai Omura, Yugo Murawaki
| Challenge: | UD Japanese resources are built on automatic conversion from several treebanks. |
| Approach: | They propose to port the word delimitation, POS, and syntactic relations of existing treebanks to UD Japanese . they discuss the issues of the UD scheme found through porting of the Japanese language . |
| Outcome: | The proposed UD Japanese resources are based on automatic conversion from treebanks. |
Constructing a Dependency Treebank for Second Language Learners of Korean (2024.lrec-main)
Copied to clipboard
| Challenge: | a manually annotated syntactic treebank is available for second language learners . the dataset includes 7,530 sentences (66,982 words; 129,333 morphemes) |
| Approach: | They propose to manually annotate syntactic treebanks based on Universal Dependencies from Korean written data. |
| Outcome: | The proposed dataset includes 7,530 sentences and 129,333 morphemes from Korean learners. |
Universal Dependencies According to BERT: Both More Specific and More General (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies show that individual BERT heads encode particular dependency relation types, but they do not match one-to-one. |
| Approach: | They propose a method for relation identification and syntactic tree construction that can be applied with minimal supervision and generalizes well across languages. |
| Outcome: | The proposed method produces significantly more consistent dependency trees than previous work and can be applied with only a minimal amount of supervision and generalizes well across languages. |
Semi-automatically Annotated Learner Corpus for Russian (2022.lrec-1)
Copied to clipboard
| Challenge: | Revita Learner Corpus is a semi-automatically annotated learner corpus for Russian . it is used for research in second language acquisition and foreign language teaching . |
| Approach: | They propose a semi-automatically annotated learner corpus for Russian that detects errors automatically and annotates errors by type. |
| Outcome: | The proposed corpus detects errors automatically and is annotated by type . the data is made public and the process is much cheaper and faster . |
Parsing Tweets into Universal Dependencies (N18-1)
Copied to clipboard
| Challenge: | a new tweet treebank for English is designed to analyze tweets with universal dependencies (UD). |
| Approach: | They extend the universal dependencies guidelines to include special constructions in tweets that affect tokenization, part-of-speech tagging, and labeled dependencies. |
| Outcome: | The proposed method outperforms state-of-the-art parsers on other treebanks in accuracy and speed. |
I Speak for the Árboles: Developing a Dependency Treebank for Spanish L2 and Heritage Speakers (2025.acl-srw)
Copied to clipboard
| Challenge: | Existing dependency treebanks for learner writing are limited due to morphosyntactic features. |
| Approach: | They propose to use a dependency treebank for Spanish learner writing from the UC Davis COWSL2H corpus to incorporate lemmatization, POS tagging, and syntactic dependencies. |
| Outcome: | The proposed treebanks are openly accessible to motivate future development of learner-oriented language technologies. |