A New Version of the Składnica Treebank of Polish Harmonised with the Walenty Valency Dictionary (L18-1)
Copied to clipboard
| Challenge: | wigra parser was used to generate the treebank, but the differences between the resources made it necessary to manually correct some parse trees. |
| Approach: | They propose a procedure to update manually disambiguated trees of Skadnica due to the switch to the Walenty valency dictionary. |
| Outcome: | The proposed method allows to check the consistency of the treebank and valence dictionary. |
Similar Papers
Aligning the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs (2022.lrec-1)
Copied to clipboard
| Challenge: | Among the language resources for Romanian, there are ones that describe the syntactic and semantic aspects of the language. |
| Approach: | They propose to align two language resources for Romanian: the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs. |
| Outcome: | The proposed alignments identify morpho-syntactic annotation mistakes, incomplete valence frames or missing ones. |
Training a Swedish Constituency Parser on Six Incompatible Treebanks (2020.lrec-1)
Copied to clipboard
| Challenge: | Syntactic parsing is a widely used intermediate step in several natural language processing tasks. |
| Approach: | They propose to use a function-tagged constituent treebank for Swedish which includes discontinuous constituents to improve the accuracy. |
| Outcome: | The proposed parser can be trained on additional treebanks that use other annotation models. |
Towards the Conversion of National Corpus of Polish to Universal Dependencies (2020.lrec-1)
Copied to clipboard
| Challenge: | a paper aims at enriching the manually annotated part of National Corpus of Polish with a syntactic layer. |
| Approach: | They enrich manually annotated part of Polish National Corpus with a syntactic layer and a UD dependency graph. |
| Outcome: | The proposed model outperforms a model trained on a smaller set of gold-standard trees in predicting part-of-speech tags, morphological features, lemmata and labelled dependency trees. |
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)
Copied to clipboard
| Challenge: | a new method is proposed to acquire typological evidence from "gold" treebanks for different languages. |
| Approach: | They propose a method for acquiring typological evidence from "gold" treebanks for different languages. |
| Outcome: | The proposed method can shed light on key issues of the linguistic typological literature. |
Parser Training with Heterogeneous Treebanks (P18-2)
Copied to clipboard
| Challenge: | In the 2017 CoNLL Shared Task on Universal Dependency Parsing, 25 languages have more than one treebank . many teams did not take advantage of the multiple treebanks, however, and trained one model per treebank instead of one model for each language. |
| Approach: | They propose a method to make the most of heterogeneous treebanks when training a monolingual parser. |
| Outcome: | The proposed method improves on training with multiple treebanks for a single language. |
NomVallex: A Valency Lexicon of Czech Nouns and Adjectives (2022.lrec-1)
Copied to clipboard
| Challenge: | NomVallex is a manual annotated valency lexicon of Czech nouns and adjectives . valencies are the ability of a verb to combine with other sentence constituents based on their morphemic forms . |
| Approach: | They propose a manually annotated valency lexicon of Czech nouns and adjectives . they capture valencies of a lexical unit in a sequence of valence slots . |
| Outcome: | The proposed lexicon is based on corpus data and contains 1027 lexical units . valency properties of lexicals are captured in a valence frame, with morphemic forms . |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
Cheating a Parser to Death: Data-driven Cross-Treebank Annotation Transfer (L18-1)
Copied to clipboard
| Challenge: | Using annotated corpus for linguistic purposes is no longer justified . hand-crafted syntactic resources such as grammars and lexicons can be used as sources of features to guide data driven systems. |
| Approach: | They propose an efficient method for transferring annotations between two different treebanks of the same language. |
| Outcome: | The proposed method is based on the Universal Dependency annotation scheme and was evaluated on the gold standard (94.75% of LAS, 99.40% UAS on the test set). |
Annotation Interoperability for the Post-ISOCat Era (2020.lrec-1)
Copied to clipboard
| Challenge: | Using ISOCat successor solutions, annotation standards have been developed since 2010 . |
| Approach: | They describe ISOCat successor solutions and annotation standardization efforts since 2010 . they describe low-cost harmonization of post-ISOCat vocabularies by means of linked ontologies . |
| Outcome: | The proposed ontologies are linked with the Ontologie of Linguistic Annotation and ISOCat, the GOLD ontology, the Typological Database Systems ontological and a large number of annotation schemes. |
A Diachronic Treebank of Russian Spanning More Than a Thousand Years (2020.lrec-1)
Copied to clipboard
| Challenge: | TOROT is a treebank that spans from the earliest Old Church Slavonic to modern Russian texts. |
| Approach: | They describe a new version of the Troms Old Russian and Old Church Slavonic Treebank . it adds a modern subcorpus to the existing treebank of contemporary standard Russian . they describe the conversion of SynTagRus into a treebank covering every attested stage of Russian and OCS . |
| Outcome: | The TOROT 20200116 treebank covers all attested stages of Russian and OCS . it includes a modern subcorpus that was created by a conversion of the SynTagRus treebank . |