| Challenge: | Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew. |
| Approach: | They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax. |
| Outcome: | The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax. |
Similar Papers
Producing a Parallel Universal Dependencies Treebank of Ancient Hebrew and Ancient Greek via Cross-Lingual Projection (2024.lrec-main)
Copied to clipboard
| Challenge: | Using parallel treebanks, syntactic changes can be identified and evaluated in translations, redactions, and commentaries. |
| Approach: | They propose to construct a treebank of Ancient Greek containing portions of the Septuagint by word-aligning and projecting from the parallel Ancient Hebrew text. |
| Outcome: | The proposed treebank contains portions of the Hebrew Scriptures, which are translated into Ancient Greek, and is based on the results of a collaborative effort to create a crosslinguistically consistent treebank annotation scheme. |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
Universal Dependencies for Amharic (L18-1)
Copied to clipboard
| Challenge: | Amharic is a morphologically rich language with a dependency relation between orthographic words and lexical categories. |
| Approach: | They propose to create an Amharic Dependency Treebank by POS tagging, morphological information and dependency relations. |
| Outcome: | The proposed treebanks are based on 1,096 sentences and are able to parse Amharic. |
A New Latin Treebank for Universal Dependencies: Charters between Ancient Latin and Romance Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, Latin features the most data and the most treebanks of all the ancient languages of UD . |
| Approach: | They introduce a Latin treebank that follows the Universal Dependencies (UD) annotation standard . they use a translation of the late Latin Charter Treebank 2 (LLCT2) into the UD style . |
| Outcome: | The proposed treebank is based on the Universal Dependencies (UD) annotation standard. |
Development of a Multilingual CCG Treebank via Universal Dependencies Conversion (2022.lrec-1)
Copied to clipboard
| Challenge: | Combinatory Categorial Grammar (CCG) is a lexicalized grammar formalism that can capture both syntactic and semantic information. |
| Approach: | They propose an algorithm to convert UD treebanks to CCG treebank and propose future extensions. |
| Outcome: | The proposed algorithm performs lexical, sentential, and syntactic rule coverage analysis, as well as CCG parsing experiments. |
Yorùbá Dependency Treebank (YTB) (2020.lrec-1)
Copied to clipboard
| Challenge: | Low-resource languages present enormous NLP opportunities as well as varying degrees of difficulties. |
| Approach: | They propose to use the Yoruba Bible treebank to apply a new grammar formalism to the language by examining the use of universal dependency annotations. |
| Outcome: | The treebank of hand-annotated parts of the Yoruba Bible provides an avenue for dependency analysis of the language; the application of a new grammar formalism to the language. |
A Second Wave of UD Hebrew Treebanking and Cross-Domain Parsing (2022.emnlp-main)
Copied to clipboard
| Challenge: | Foundational Hebrew NLP tasks have relied on various versions of the Hebrew Treebank . however, the data in the HTB is now over 30 years old and does not cover many aspects of contemporary Hebrew on the web. |
| Approach: | They propose to use Hebrew Wikipedia to stratify the text from a UD treebank. |
| Outcome: | The proposed treebank is based on a single-source newswire corpus selected from Hebrew Wikipedia. |
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)
Copied to clipboard
| Challenge: | spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all. |
| Approach: | They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme. |
| Outcome: | The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation. |
Building Universal Dependency Treebanks in Korean (L18-1)
Copied to clipboard
| Challenge: | Several treebanks were introduced for Korean, all of which comprised annotation of morphemes and phrase structure trees, each following its own set of guidelines. |
| Approach: | They propose to use Korean treebanks as dependency trees and to analyze their performance using morpheme-level annotations. |
| Outcome: | The Korean portion of the Google UD Treebank, the Penn Korean Treebank and the KAIST Treebank are re-tokenized and assessed for errors. |
Creating a Parallel Icelandic Dependency Treebank from Raw Text to Universal Dependencies (2020.lrec-1)
Copied to clipboard
| Challenge: | Icelandic language is low-resource and is not yet considered in imminent danger . efforts underway to make it accessible and usable in Language Technology . |
| Approach: | They propose to build a parallel Icelandic dependency treebank based on Universal Dependencies (UD) this is the first parallel treebank resource for the language and several other languages already have one . |
| Outcome: | The proposed treebank is the first parallel treebank resource for the low-resource language . the project will be published as part of UD version 2.6. |