Investigating UD Treebanks via Dataset Difficulty Measures (2023.eacl-main)

Copied to clipboard

Challenge: Treebanks annotated with Universal Dependencies (UD) are currently available for over 100 languages and are only partially reflected in parser evaluations via accuracy metrics like LAS.
Approach: They propose to use dataset cartography, V-information, and minimum description length to analyze UD treebanks using three accuracy-free methods to provide insights about them.
Outcome: The proposed methods provide insights about UD treebanks that would remain undetected if only LAS was considered.

Similar Papers

Development of a Multilingual CCG Treebank via Universal Dependencies Conversion (2022.lrec-1)

Copied to clipboard

Challenge: Combinatory Categorial Grammar (CCG) is a lexicalized grammar formalism that can capture both syntactic and semantic information.
Approach: They propose an algorithm to convert UD treebanks to CCG treebank and propose future extensions.
Outcome: The proposed algorithm performs lexical, sentential, and syntactic rule coverage analysis, as well as CCG parsing experiments.
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)

Copied to clipboard

Challenge: Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages.
Approach: They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer .
Outcome: The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers.
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)

Copied to clipboard

Challenge: spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all.
Approach: They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme.
Outcome: The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation.
Parsing Tweets into Universal Dependencies (N18-1)

Copied to clipboard

Challenge: a new tweet treebank for English is designed to analyze tweets with universal dependencies (UD).
Approach: They extend the universal dependencies guidelines to include special constructions in tweets that affect tokenization, part-of-speech tagging, and labeled dependencies.
Outcome: The proposed method outperforms state-of-the-art parsers on other treebanks in accuracy and speed.
Scalable Cross-lingual Treebank Synthesis for Improved Production Dependency Parsers (2020.coling-industry)

Copied to clipboard

Challenge: scalable Universal Dependency (UD) treebank synthesis techniques are used to improve production-grade parsers.
Approach: They propose a data augmentation technique that uses synthetic treebanks to improve production-grade parsers.
Outcome: The proposed technique improves LAS performance on seven languages by up to two points on production models trained on original UD treebanks.
Some Languages Seem Easier to Parse Because Their Treebanks Leak (2020.emnlp-main)

Copied to clipboard

Challenge: Cross-language differences in (universal) dependency parsing performance are mostly attributed to treebank size, average sentence length, average dependency length, morphological complexity, and domain differences.
Approach: They compute graph isomorphisms and find that treebank size is a factor that influences parsing performance.
Outcome: The results show that the overlap between training and test graphs explain more of the observed variation than standard explanations such as the above.
Building Universal Dependency Treebanks in Korean (L18-1)

Copied to clipboard

Challenge: Several treebanks were introduced for Korean, all of which comprised annotation of morphemes and phrase structure trees, each following its own set of guidelines.
Approach: They propose to use Korean treebanks as dependency trees and to analyze their performance using morpheme-level annotations.
Outcome: The Korean portion of the Google UD Treebank, the Penn Korean Treebank and the KAIST Treebank are re-tokenized and assessed for errors.
A Universal Dependencies Treebank of Ancient Hebrew (2022.lrec-1)

Copied to clipboard

Challenge: Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew.
Approach: They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax.
Outcome: The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax.
Universal Dependencies Version 2 for Japanese (L18-1)

Copied to clipboard

Challenge: UD Japanese resources are built on automatic conversion from several treebanks.
Approach: They propose to port the word delimitation, POS, and syntactic relations of existing treebanks to UD Japanese . they discuss the issues of the UD scheme found through porting of the Japanese language .
Outcome: The proposed UD Japanese resources are based on automatic conversion from treebanks.
The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation Project (2025.acl-long)

Copied to clipboard

Challenge: UD-NewsCrawl is the largest Tagalog treebank to date, with 15.6k trees manually annotated according to the Universal Dependencies framework.
Approach: They propose to use UD-NewsCrawl to annotate Tagalog trees using the Universal Dependencies framework.
Outcome: The proposed treebanks are based on the Universal Dependencies framework and have 15.6k trees annotated manually.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations