Challenge: spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all.
Approach: They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme.
Outcome: The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation.

Similar Papers

Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)

Copied to clipboard

Challenge: Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages.
Approach: They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer .
Outcome: The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers.
UCxn: Typologically-Informed Annotation of Constructions Atop Universal Dependencies (2024.lrec-main)

Copied to clipboard

Challenge: Grammatical constructions that convey meaning through a particular combination of several morphosyntactic elements are not labeled holistically.
Approach: They propose to augment UD annotations with a ‘UCxn’ annotation layer for such meaning-bearing grammatical constructions and to approach this in a typologically informed way so that morphosyntactic strategies can be compared across languages.
Outcome: The proposed annotation layer could be used to annotate meaning-bearing constructions across languages and to compare them across languages.
Building Universal Dependency Treebanks in Korean (L18-1)

Copied to clipboard

Challenge: Several treebanks were introduced for Korean, all of which comprised annotation of morphemes and phrase structure trees, each following its own set of guidelines.
Approach: They propose to use Korean treebanks as dependency trees and to analyze their performance using morpheme-level annotations.
Outcome: The Korean portion of the Google UD Treebank, the Penn Korean Treebank and the KAIST Treebank are re-tokenized and assessed for errors.
Investigating UD Treebanks via Dataset Difficulty Measures (2023.eacl-main)

Copied to clipboard

Challenge: Treebanks annotated with Universal Dependencies (UD) are currently available for over 100 languages and are only partially reflected in parser evaluations via accuracy metrics like LAS.
Approach: They propose to use dataset cartography, V-information, and minimum description length to analyze UD treebanks using three accuracy-free methods to provide insights about them.
Outcome: The proposed methods provide insights about UD treebanks that would remain undetected if only LAS was considered.
Development of a Multilingual CCG Treebank via Universal Dependencies Conversion (2022.lrec-1)

Copied to clipboard

Challenge: Combinatory Categorial Grammar (CCG) is a lexicalized grammar formalism that can capture both syntactic and semantic information.
Approach: They propose an algorithm to convert UD treebanks to CCG treebank and propose future extensions.
Outcome: The proposed algorithm performs lexical, sentential, and syntactic rule coverage analysis, as well as CCG parsing experiments.
Treebanking User-Generated Content: A Proposal for a Unified Representation in Universal Dependencies (2020.lrec-1)

Copied to clipboard

Challenge: Despite the increasing number of contributions on Part-of-Speech tagging and parsing, automatic processing of user-generated content (UGC) still represents a challenging task.
Approach: They propose a set of guidelines for the annotation of user-generated texts within the Universal Dependencies framework.
Outcome: The proposed annotation guidelines promote cross-linguistic consistency, which has always been in the spirit of UD.
MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank (2024.lrec-main)

Copied to clipboard

Challenge: Despite the success of the Universal Dependencies (UD) project, there is still a lack of diversity within high-resource languages and their closely related non-standard languages and dialects.
Approach: They propose to annotate Bavarian with part-of-speech and syntactic dependency information manually in UD and to highlight morphosyntactical differences between the closely related languages.
Outcome: The proposed treebank covers multiple genres including wiki, fiction, grammar examples, social, non-fiction and Bavarian.
A Universal Dependencies Treebank of Ancient Hebrew (2022.lrec-1)

Copied to clipboard

Challenge: Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew.
Approach: They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax.
Outcome: The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax.
I Speak for the Árboles: Developing a Dependency Treebank for Spanish L2 and Heritage Speakers (2025.acl-srw)

Copied to clipboard

Challenge: Existing dependency treebanks for learner writing are limited due to morphosyntactic features.
Approach: They propose to use a dependency treebank for Spanish learner writing from the UC Davis COWSL2H corpus to incorporate lemmatization, POS tagging, and syntactic dependencies.
Outcome: The proposed treebanks are openly accessible to motivate future development of learner-oriented language technologies.
A New Latin Treebank for Universal Dependencies: Charters between Ancient Latin and Romance Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, Latin features the most data and the most treebanks of all the ancient languages of UD .
Approach: They introduce a Latin treebank that follows the Universal Dependencies (UD) annotation standard . they use a translation of the late Latin Charter Treebank 2 (LLCT2) into the UD style .
Outcome: The proposed treebank is based on the Universal Dependencies (UD) annotation standard.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations