Jan Odijk, Alexis Dimitriadis, Martijn van der Klis, Marjo van Koppen, Meie Otten, Remco van der Veen
| Challenge: | Using the AnnCor CHILDES Treebank, we assign adult grammar syntactic structures to children's utterances. |
| Approach: | They propose a partially manually verified treebank for Dutch CHILDES corpora . they argue that human annotation and automatic checks on this annotation must go hand in hand . |
| Outcome: | The AnnCor CHILDES Treebank is the first partially manually verified treebank for Dutch CHILdes corpora. |
Similar Papers
Establishing a New State-of-the-Art for French Named Entity Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a task consisting in identifying text spans that denote named entities such as person, location and organization names. |
| Approach: | They manually annotated the French TreeBank with information related to named entities . they sketch the underlying annotation guidelines and provide a few figures about the annotations . |
| Outcome: | The French TreeBank is the main source of morphosyntactic and syntactical annotations for French. |
A Universal Dependencies Treebank of Ancient Hebrew (2022.lrec-1)
Copied to clipboard
| Challenge: | Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew. |
| Approach: | They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax. |
| Outcome: | The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax. |
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)
Copied to clipboard
| Challenge: | spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all. |
| Approach: | They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme. |
| Outcome: | The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation. |
Building a Universal Dependencies Treebank for Occitan (2020.lrec-1)
Copied to clipboard
Aleksandra Miletic, Myriam Bras, Marianne Vergez-Couret, Louise Esher, Clamença Poujade, Jean Sibille
| Challenge: | Low-resourced regional, non-official or minority languages often face lack of institutional support . low-resource languages often find themselves in a similar situation . |
| Approach: | They propose to create the first treebank for Occitan, a low-resourced regional language . they use an agile annotation approach and rely on pre-processing using existing tools . |
| Outcome: | The proposed treebank is the first for the low-resourced regional language Occitan . the project uses an agile annotation approach and automated pre-annotation . |
The Tembusu Treebank: An English Learner Treebank (2022.lrec-1)
Copied to clipboard
| Challenge: | a new treebank is created to help diagnose ungrammatical sentences using mal-rules . the Tembusu Learner Treebank is an open treebank created from the corpus of Learner English . |
| Approach: | They propose to use the Tembusu Learner Treebank to train a new parse-ranking model for the English Resource Grammar . the model incorporates mal-rules in the annotation of ungrammatical sentences . |
| Outcome: | The Tembusu Learner Treebank is an open treebank created from the NTU Corpus of Learner English . the treebank is unique for incorporating mal-rules in the annotation of ungrammatical sentences . |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
A Penn-style Treebank of Middle Low German (2020.lrec-1)
Copied to clipboard
| Challenge: | attestation for Middle Low German is rich, but its syntax remains relatively understudied. |
| Approach: | They outline the issues involved in creating a Penn-style treebank of Middle Low German . they describe the background for the corpus and the process by which texts were selected . |
| Outcome: | The proposed corpus will be a syntactically annotated treebank of Middle Low German . the proposed corpuse will be part of the Corpus of Historical Low German (CHLG) the proposed method will be used to generate strong empirical evidence for the language . |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
Cheating a Parser to Death: Data-driven Cross-Treebank Annotation Transfer (L18-1)
Copied to clipboard
| Challenge: | Using annotated corpus for linguistic purposes is no longer justified . hand-crafted syntactic resources such as grammars and lexicons can be used as sources of features to guide data driven systems. |
| Approach: | They propose an efficient method for transferring annotations between two different treebanks of the same language. |
| Outcome: | The proposed method is based on the Universal Dependency annotation scheme and was evaluated on the gold standard (94.75% of LAS, 99.40% UAS on the test set). |
Yorùbá Dependency Treebank (YTB) (2020.lrec-1)
Copied to clipboard
| Challenge: | Low-resource languages present enormous NLP opportunities as well as varying degrees of difficulties. |
| Approach: | They propose to use the Yoruba Bible treebank to apply a new grammar formalism to the language by examining the use of universal dependency annotations. |
| Outcome: | The treebank of hand-annotated parts of the Yoruba Bible provides an avenue for dependency analysis of the language; the application of a new grammar formalism to the language. |