| Challenge: | ellipsis is a common linguistic phenomenon that some words are left out as they are understood from the context, especially in oral utterance. |
| Approach: | They propose to use a Chinese dependency treebank to facilitate the parsing of web text . they propose to restore omissions and reserve contexts in the web text to improve dependency parsers . |
| Outcome: | The proposed framework enables the parsing of web text from online microblogs. |
Similar Papers
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
Semantic-pragmatic Annotations in the Prague Dependency Treebank (2026.findings-acl)
Copied to clipboard
Marie Mikulová, Eva Hajicova, Jiří Mírovský, Anna Nedoluzhko, Michal Novák, Pavlína Synková, Jan Štěpánek, Barbora Štěpánková, Jan Hajič
| Challenge: | morphology and syntax work on sentence level, but semantic-pragmatic phenomena are often related to two or more neighbouring sentences and possibly to an extra-linguistic context. |
| Approach: | They present semantic-pragmatic specification and annotations in the Prague Dependency Treebank - Consolidated 2.0 release2 by annotating the entire corpus. |
| Outcome: | The proposed annotations are based on the Prague Dependency Treebank -Consolidated 2.0 (PDT-C 2.0) the dataset contains more than 3 million tokens (of Czech) manually annotated from morphology to surface and deep syntax including several types of semantic-pragmatic annotations. |
The Persian Dependency Treebank Made Universal (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing universal dependency treebanks are lacking sufficient annotated data. |
| Approach: | They propose a method for converting Persian Dependency Treebank to Universal Dependencies using an automatic method. |
| Outcome: | The proposed method is more compatible with Universal Dependencies than the Uppsala Persian Universal Dependency Treebank. |
L1-L2 Parallel Treebank of Learner Chinese: Overused and Underused Syntactic Structures (L18-1)
Copied to clipboard
| Challenge: | Currently, the treebank consists of 600 L2 sentences and 697 L1 sentences. |
| Approach: | They propose to use "L1-L2 parallel treebanks" to facilitate analyses of learner language. |
| Outcome: | The proposed treebank consists of 600 L2 sentences and 697 L1 sentences. |
BKTreebank: Building a Vietnamese Dependency Treebank (L18-1)
Copied to clipboard
| Challenge: | In this paper, we present the building of a dependency treebank for Vietnamese . |
| Approach: | They propose to build a Vietnamese dependency treebank using automatic taggers and automatic tagging. |
| Outcome: | The proposed treebank is a useful resource for Vietnamese language processing. |
Treebank Embedding Vectors for Out-of-Domain Dependency Parsing (2020.acl-main)
Copied to clipboard
| Challenge: | a recent advance in monolingual dependency parsing is the idea of a treebank embedding vector . this allows the model to prefer training data from one treebank over another at test time . |
| Approach: | They propose a method to predict a treebank vector for sentences that do not come from a particular treebank . they also explore what happens when they move away from predefined treebank embedding vectors . |
| Outcome: | The proposed method can predict treebank vectors for sentences that do not come from a treebank used in training with sufficient accuracy for nine out of ten languages. |
Parser Training with Heterogeneous Treebanks (P18-2)
Copied to clipboard
| Challenge: | In the 2017 CoNLL Shared Task on Universal Dependency Parsing, 25 languages have more than one treebank . many teams did not take advantage of the multiple treebanks, however, and trained one model per treebank instead of one model for each language. |
| Approach: | They propose a method to make the most of heterogeneous treebanks when training a monolingual parser. |
| Outcome: | The proposed method improves on training with multiple treebanks for a single language. |
Camel Treebank: An Open Multi-genre Arabic Dependency Treebank (2022.lrec-1)
Copied to clipboard
| Challenge: | CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses) |
| Approach: | They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic. |
| Outcome: | The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses) |
A Gold Standard Dependency Treebank for Turkish (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, Turkish treebanks are limited due to the limited number of annotated sentences in the domains of Wikipedia and ITU Web Treebanks. |
| Approach: | They propose to annotate Turkish web and Wikipedia sentences for segmentation, morphology, part-of-speech and dependency relations using tagsets and a Wikipedia section. |
| Outcome: | The proposed treebank is the largest publicly available morpho-syntactic treebank in terms of word count and has a dedicated Wikipedia section. |
Constructing a Dependency Treebank for Second Language Learners of Korean (2024.lrec-main)
Copied to clipboard
| Challenge: | a manually annotated syntactic treebank is available for second language learners . the dataset includes 7,530 sentences (66,982 words; 129,333 morphemes) |
| Approach: | They propose to manually annotate syntactic treebanks based on Universal Dependencies from Korean written data. |
| Outcome: | The proposed dataset includes 7,530 sentences and 129,333 morphemes from Korean learners. |