Challenge: wigra parser was used to generate the treebank, but the differences between the resources made it necessary to manually correct some parse trees.
Approach: They propose a procedure to update manually disambiguated trees of Skadnica due to the switch to the Walenty valency dictionary.
Outcome: The proposed method allows to check the consistency of the treebank and valence dictionary.

Similar Papers

Aligning the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs (2022.lrec-1)

Copied to clipboard

Challenge: Among the language resources for Romanian, there are ones that describe the syntactic and semantic aspects of the language.
Approach: They propose to align two language resources for Romanian: the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs.
Outcome: The proposed alignments identify morpho-syntactic annotation mistakes, incomplete valence frames or missing ones.
Training a Swedish Constituency Parser on Six Incompatible Treebanks (2020.lrec-1)

Copied to clipboard

Challenge: Syntactic parsing is a widely used intermediate step in several natural language processing tasks.
Approach: They propose to use a function-tagged constituent treebank for Swedish which includes discontinuous constituents to improve the accuracy.
Outcome: The proposed parser can be trained on additional treebanks that use other annotation models.
Towards the Conversion of National Corpus of Polish to Universal Dependencies (2020.lrec-1)

Copied to clipboard

Challenge: a paper aims at enriching the manually annotated part of National Corpus of Polish with a syntactic layer.
Approach: They enrich manually annotated part of Polish National Corpus with a syntactic layer and a UD dependency graph.
Outcome: The proposed model outperforms a model trained on a smaller set of gold-standard trees in predicting part-of-speech tags, morphological features, lemmata and labelled dependency trees.
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)

Copied to clipboard

Challenge: a new method is proposed to acquire typological evidence from "gold" treebanks for different languages.
Approach: They propose a method for acquiring typological evidence from "gold" treebanks for different languages.
Outcome: The proposed method can shed light on key issues of the linguistic typological literature.
Parser Training with Heterogeneous Treebanks (P18-2)

Copied to clipboard

Challenge: In the 2017 CoNLL Shared Task on Universal Dependency Parsing, 25 languages have more than one treebank . many teams did not take advantage of the multiple treebanks, however, and trained one model per treebank instead of one model for each language.
Approach: They propose a method to make the most of heterogeneous treebanks when training a monolingual parser.
Outcome: The proposed method improves on training with multiple treebanks for a single language.
NomVallex: A Valency Lexicon of Czech Nouns and Adjectives (2022.lrec-1)

Copied to clipboard

Challenge: NomVallex is a manual annotated valency lexicon of Czech nouns and adjectives . valencies are the ability of a verb to combine with other sentence constituents based on their morphemic forms .
Approach: They propose a manually annotated valency lexicon of Czech nouns and adjectives . they capture valencies of a lexical unit in a sequence of valence slots .
Outcome: The proposed lexicon is based on corpus data and contains 1027 lexical units . valency properties of lexicals are captured in a valence frame, with morphemic forms .
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)

Copied to clipboard

Challenge: Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages.
Approach: They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer .
Outcome: The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers.
Cheating a Parser to Death: Data-driven Cross-Treebank Annotation Transfer (L18-1)

Copied to clipboard

Challenge: Using annotated corpus for linguistic purposes is no longer justified . hand-crafted syntactic resources such as grammars and lexicons can be used as sources of features to guide data driven systems.
Approach: They propose an efficient method for transferring annotations between two different treebanks of the same language.
Outcome: The proposed method is based on the Universal Dependency annotation scheme and was evaluated on the gold standard (94.75% of LAS, 99.40% UAS on the test set).
Annotation Interoperability for the Post-ISOCat Era (2020.lrec-1)

Copied to clipboard

Challenge: Using ISOCat successor solutions, annotation standards have been developed since 2010 .
Approach: They describe ISOCat successor solutions and annotation standardization efforts since 2010 . they describe low-cost harmonization of post-ISOCat vocabularies by means of linked ontologies .
Outcome: The proposed ontologies are linked with the Ontologie of Linguistic Annotation and ISOCat, the GOLD ontology, the Typological Database Systems ontological and a large number of annotation schemes.
A Diachronic Treebank of Russian Spanning More Than a Thousand Years (2020.lrec-1)

Copied to clipboard

Challenge: TOROT is a treebank that spans from the earliest Old Church Slavonic to modern Russian texts.
Approach: They describe a new version of the Troms Old Russian and Old Church Slavonic Treebank . it adds a modern subcorpus to the existing treebank of contemporary standard Russian . they describe the conversion of SynTagRus into a treebank covering every attested stage of Russian and OCS .
Outcome: The TOROT 20200116 treebank covers all attested stages of Russian and OCS . it includes a modern subcorpus that was created by a conversion of the SynTagRus treebank .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations