ForFun 1.0: Prague Database of Forms and Functions – An Invaluable Resource for Linguistic Research (L18-1)
Copied to clipboard
| Challenge: | a new database, ForFun, is an invaluable resource for linguistic research . it allows researchers to elaborate various syntactic issues in depth . |
| Approach: | They introduce the first version of ForFun, Prague Database of Forms and Functions . it uses syntactically annotated Prague Dependency Treebanks to organize annotations anew . |
| Outcome: | The proposed database brings syntactic issues closer to researchers . it takes advantage of existing treebanks and offers a user-friendly access to real examples . |
Similar Papers
Prague Dependency Treebank - Consolidated 1.0 (2020.lrec-1)
Copied to clipboard
Jan Hajič, Eduard Bejček, Jaroslava Hlavacova, Marie Mikulová, Milan Straka, Jan Štěpánek, Barbora Štěpánková
| Challenge: | Using the standard PDT scheme, the Prague Dependency Treebank-Consolidated 1.0 contains 4 different datasets of Czech, uniformly annotated using the standard scheme. |
| Approach: | They present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0, which contains 4 different datasets of Czech, uniformly annnotated using the standard PDT scheme. |
| Outcome: | The Prague Dependency Treebank-Consolidated 1.0 contains 4 datasets of Czech, uniformly annotated using the standard PDT scheme. |
From Form to Meaning: The Case of Particles within the Prague Dependency Treebank Annotation Scheme (2025.coling-main)
Copied to clipboard
| Challenge: | Discussions on an appropriate annotation scheme for large and complex information are ongoing . multi-layer system allows a comprehensive description of relations between morphological properties, syntactic function and expressed meaning. |
| Approach: | They propose a multi-layer annotation scheme for the Prague Dependency Treebank . they propose morphological properties, syntactic function and expressed meaning as multi-layered systems . |
| Outcome: | The proposed scheme is sound and serves well for complex annotations. |
Semantic-pragmatic Annotations in the Prague Dependency Treebank (2026.findings-acl)
Copied to clipboard
Marie Mikulová, Eva Hajicova, Jiří Mírovský, Anna Nedoluzhko, Michal Novák, Pavlína Synková, Jan Štěpánek, Barbora Štěpánková, Jan Hajič
| Challenge: | morphology and syntax work on sentence level, but semantic-pragmatic phenomena are often related to two or more neighbouring sentences and possibly to an extra-linguistic context. |
| Approach: | They present semantic-pragmatic specification and annotations in the Prague Dependency Treebank - Consolidated 2.0 release2 by annotating the entire corpus. |
| Outcome: | The proposed annotations are based on the Prague Dependency Treebank -Consolidated 2.0 (PDT-C 2.0) the dataset contains more than 3 million tokens (of Czech) manually annotated from morphology to surface and deep syntax including several types of semantic-pragmatic annotations. |
Cost-Effective Discourse Annotation in the Prague Czech–English Dependency Treebank (2024.lrec-main)
Copied to clipboard
| Challenge: | a method for obtaining a high-quality annotation of explicit discourse relations is a resource-demanding task. |
| Approach: | They propose a method for obtaining a high-quality annotation of explicit discourse relations in the Czech part of the Prague Czech–English Dependency Treebank. |
| Outcome: | The proposed method solves the problem of identifying discrepancies between the annotations in the Czech part of the Penn Treebank. |
Announcing the Prague Discourse Treebank 3.0 (2024.lrec-main)
Copied to clipboard
| Challenge: | PDiT 3.0 contains 21,662 discourse relations (plus 445 list relations) in 49 thousand sentences. |
| Approach: | They present the Prague Discourse Treebank 3.0, a new version of the annotation of discourse relations marked by primary and secondary discourse connectives in the Prague Dependency Treebank. |
| Outcome: | The new version of the PDiT 3.0 brings a largely revised annotation of discourse relations and achieves consistency with a Lexicon of Czech Discourse Connectives (CzeDLex) and sense taxonomy. |
MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank (2024.lrec-main)
Copied to clipboard
| Challenge: | Despite the success of the Universal Dependencies (UD) project, there is still a lack of diversity within high-resource languages and their closely related non-standard languages and dialects. |
| Approach: | They propose to annotate Bavarian with part-of-speech and syntactic dependency information manually in UD and to highlight morphosyntactical differences between the closely related languages. |
| Outcome: | The proposed treebank covers multiple genres including wiki, fiction, grammar examples, social, non-fiction and Bavarian. |
Fine-grained Classification of Circumstantial Meanings within the Prague Dependency Treebank Annotation Scheme (2024.lrec-main)
Copied to clipboard
| Challenge: | a formally and semantically based fine-grained classification of circumstantial meanings is proposed for the Czech language . the methodology and principles used are language independent . |
| Approach: | They propose a formally and semantically based fine-grained classification of circumstantial meanings based on Prague Dependency Treebanks examples. |
| Outcome: | The proposed method is language independent and compares with English . it is carried out in the Czech language but not in any other annotation project . |
The Treebank of Vedic Sanskrit (2020.lrec-1)
Copied to clipboard
| Challenge: | Vedic Sanskrit is a morphologically rich ancient Indian language of central importance for linguistic and historical research. |
| Approach: | They introduce the first treebank of Vedic Sanskrit, a morphologically rich ancient Indian language . they describe how sentences are annotated in the Universal Dependencies scheme and which syntactic constructions required special attention. |
| Outcome: | The proposed treebank reflects the development of metrical and prose texts over a period of 600 years. |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
GPT-SW3: An Autoregressive Language Model for the Scandinavian Languages (2024.lrec-main)
Copied to clipboard
Ariel Ekgren, Amaru Cuba Gyllensten, Felix Stollenwerk, Joey Öhman, Tim Isbister, Evangelia Gogoulou, Fredrik Carlsson, Judit Casademont, Magnus Sahlgren
| Challenge: | a growing interest in building and applying large language models for languages other than English is fueling interest in developing LLMs for smaller languages. |
| Approach: | They describe the development process for the first native large generative language model for the North Germanic languages, GPT-SW3. |
| Outcome: | The proposed model is based on the generative language model for the North Germanic languages . it is a first-generation model with a high-quality data set and a low cost of implementation . |