Challenge: a new database, ForFun, is an invaluable resource for linguistic research . it allows researchers to elaborate various syntactic issues in depth .
Approach: They introduce the first version of ForFun, Prague Database of Forms and Functions . it uses syntactically annotated Prague Dependency Treebanks to organize annotations anew .
Outcome: The proposed database brings syntactic issues closer to researchers . it takes advantage of existing treebanks and offers a user-friendly access to real examples .

Similar Papers

Prague Dependency Treebank - Consolidated 1.0 (2020.lrec-1)

Copied to clipboard

Challenge: Using the standard PDT scheme, the Prague Dependency Treebank-Consolidated 1.0 contains 4 different datasets of Czech, uniformly annotated using the standard scheme.
Approach: They present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0, which contains 4 different datasets of Czech, uniformly annnotated using the standard PDT scheme.
Outcome: The Prague Dependency Treebank-Consolidated 1.0 contains 4 datasets of Czech, uniformly annotated using the standard PDT scheme.
From Form to Meaning: The Case of Particles within the Prague Dependency Treebank Annotation Scheme (2025.coling-main)

Copied to clipboard

Challenge: Discussions on an appropriate annotation scheme for large and complex information are ongoing . multi-layer system allows a comprehensive description of relations between morphological properties, syntactic function and expressed meaning.
Approach: They propose a multi-layer annotation scheme for the Prague Dependency Treebank . they propose morphological properties, syntactic function and expressed meaning as multi-layered systems .
Outcome: The proposed scheme is sound and serves well for complex annotations.
Semantic-pragmatic Annotations in the Prague Dependency Treebank (2026.findings-acl)

Copied to clipboard

Challenge: morphology and syntax work on sentence level, but semantic-pragmatic phenomena are often related to two or more neighbouring sentences and possibly to an extra-linguistic context.
Approach: They present semantic-pragmatic specification and annotations in the Prague Dependency Treebank - Consolidated 2.0 release2 by annotating the entire corpus.
Outcome: The proposed annotations are based on the Prague Dependency Treebank -Consolidated 2.0 (PDT-C 2.0) the dataset contains more than 3 million tokens (of Czech) manually annotated from morphology to surface and deep syntax including several types of semantic-pragmatic annotations.
Cost-Effective Discourse Annotation in the Prague Czech–English Dependency Treebank (2024.lrec-main)

Copied to clipboard

Challenge: a method for obtaining a high-quality annotation of explicit discourse relations is a resource-demanding task.
Approach: They propose a method for obtaining a high-quality annotation of explicit discourse relations in the Czech part of the Prague Czech–English Dependency Treebank.
Outcome: The proposed method solves the problem of identifying discrepancies between the annotations in the Czech part of the Penn Treebank.
Announcing the Prague Discourse Treebank 3.0 (2024.lrec-main)

Copied to clipboard

Challenge: PDiT 3.0 contains 21,662 discourse relations (plus 445 list relations) in 49 thousand sentences.
Approach: They present the Prague Discourse Treebank 3.0, a new version of the annotation of discourse relations marked by primary and secondary discourse connectives in the Prague Dependency Treebank.
Outcome: The new version of the PDiT 3.0 brings a largely revised annotation of discourse relations and achieves consistency with a Lexicon of Czech Discourse Connectives (CzeDLex) and sense taxonomy.
MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank (2024.lrec-main)

Copied to clipboard

Challenge: Despite the success of the Universal Dependencies (UD) project, there is still a lack of diversity within high-resource languages and their closely related non-standard languages and dialects.
Approach: They propose to annotate Bavarian with part-of-speech and syntactic dependency information manually in UD and to highlight morphosyntactical differences between the closely related languages.
Outcome: The proposed treebank covers multiple genres including wiki, fiction, grammar examples, social, non-fiction and Bavarian.
Fine-grained Classification of Circumstantial Meanings within the Prague Dependency Treebank Annotation Scheme (2024.lrec-main)

Copied to clipboard

Challenge: a formally and semantically based fine-grained classification of circumstantial meanings is proposed for the Czech language . the methodology and principles used are language independent .
Approach: They propose a formally and semantically based fine-grained classification of circumstantial meanings based on Prague Dependency Treebanks examples.
Outcome: The proposed method is language independent and compares with English . it is carried out in the Czech language but not in any other annotation project .
The Treebank of Vedic Sanskrit (2020.lrec-1)

Copied to clipboard

Challenge: Vedic Sanskrit is a morphologically rich ancient Indian language of central importance for linguistic and historical research.
Approach: They introduce the first treebank of Vedic Sanskrit, a morphologically rich ancient Indian language . they describe how sentences are annotated in the Universal Dependencies scheme and which syntactic constructions required special attention.
Outcome: The proposed treebank reflects the development of metrical and prose texts over a period of 600 years.
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)

Copied to clipboard

Challenge: Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages.
Approach: They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer .
Outcome: The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers.
GPT-SW3: An Autoregressive Language Model for the Scandinavian Languages (2024.lrec-main)

Copied to clipboard

Challenge: a growing interest in building and applying large language models for languages other than English is fueling interest in developing LLMs for smaller languages.
Approach: They describe the development process for the first native large generative language model for the North Germanic languages, GPT-SW3.
Outcome: The proposed model is based on the generative language model for the North Germanic languages . it is a first-generation model with a high-quality data set and a low cost of implementation .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations