From Form to Meaning: The Case of Particles within the Prague Dependency Treebank Annotation Scheme (2025.coling-main)
Copied to clipboard
| Challenge: | Discussions on an appropriate annotation scheme for large and complex information are ongoing . multi-layer system allows a comprehensive description of relations between morphological properties, syntactic function and expressed meaning. |
| Approach: | They propose a multi-layer annotation scheme for the Prague Dependency Treebank . they propose morphological properties, syntactic function and expressed meaning as multi-layered systems . |
| Outcome: | The proposed scheme is sound and serves well for complex annotations. |
Similar Papers
Semantic-pragmatic Annotations in the Prague Dependency Treebank (2026.findings-acl)
Copied to clipboard
Marie Mikulová, Eva Hajicova, Jiří Mírovský, Anna Nedoluzhko, Michal Novák, Pavlína Synková, Jan Štěpánek, Barbora Štěpánková, Jan Hajič
| Challenge: | morphology and syntax work on sentence level, but semantic-pragmatic phenomena are often related to two or more neighbouring sentences and possibly to an extra-linguistic context. |
| Approach: | They present semantic-pragmatic specification and annotations in the Prague Dependency Treebank - Consolidated 2.0 release2 by annotating the entire corpus. |
| Outcome: | The proposed annotations are based on the Prague Dependency Treebank -Consolidated 2.0 (PDT-C 2.0) the dataset contains more than 3 million tokens (of Czech) manually annotated from morphology to surface and deep syntax including several types of semantic-pragmatic annotations. |
Fine-grained Classification of Circumstantial Meanings within the Prague Dependency Treebank Annotation Scheme (2024.lrec-main)
Copied to clipboard
| Challenge: | a formally and semantically based fine-grained classification of circumstantial meanings is proposed for the Czech language . the methodology and principles used are language independent . |
| Approach: | They propose a formally and semantically based fine-grained classification of circumstantial meanings based on Prague Dependency Treebanks examples. |
| Outcome: | The proposed method is language independent and compares with English . it is carried out in the Czech language but not in any other annotation project . |
ForFun 1.0: Prague Database of Forms and Functions – An Invaluable Resource for Linguistic Research (L18-1)
Copied to clipboard
| Challenge: | a new database, ForFun, is an invaluable resource for linguistic research . it allows researchers to elaborate various syntactic issues in depth . |
| Approach: | They introduce the first version of ForFun, Prague Database of Forms and Functions . it uses syntactically annotated Prague Dependency Treebanks to organize annotations anew . |
| Outcome: | The proposed database brings syntactic issues closer to researchers . it takes advantage of existing treebanks and offers a user-friendly access to real examples . |
A Survey of Meaning Representations – From Theory to Practical Utility (2024.naacl-long)
Copied to clipboard
| Challenge: | Symbolic meaning representations of natural language text have been studied since at least the 1960s . with the availability of large annotated corpora, the field has recently seen several new developments . |
| Approach: | They propose a framework for expressing meaning in natural language text using annotated corpora and a set of tools for machine learning. |
| Outcome: | The frameworks are based on a set of theoretical and practical problems and their applications. |
Joint Annotation of Morphology and Syntax in Dependency Treebanks (2024.lrec-main)
Copied to clipboard
| Challenge: | Syntactic treebanks have been in development since the 1970s . they are now available for a vast array of languages from across the globe . |
| Approach: | They propose new formats to annotate syntactic and morphological relations in a dependency treebank using distributional criteria for the choice of the head of any combination. |
| Outcome: | The proposed formats are compatible with the UD schema for syntactic treebanks. |
Prague Dependency Treebank - Consolidated 1.0 (2020.lrec-1)
Copied to clipboard
Jan Hajič, Eduard Bejček, Jaroslava Hlavacova, Marie Mikulová, Milan Straka, Jan Štěpánek, Barbora Štěpánková
| Challenge: | Using the standard PDT scheme, the Prague Dependency Treebank-Consolidated 1.0 contains 4 different datasets of Czech, uniformly annotated using the standard scheme. |
| Approach: | They present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0, which contains 4 different datasets of Czech, uniformly annnotated using the standard PDT scheme. |
| Outcome: | The Prague Dependency Treebank-Consolidated 1.0 contains 4 datasets of Czech, uniformly annotated using the standard PDT scheme. |
Treebanking User-Generated Content: A Proposal for a Unified Representation in Universal Dependencies (2020.lrec-1)
Copied to clipboard
Manuela Sanguinetti, Cristina Bosco, Lauren Cassidy, Özlem Çetinoğlu, Alessandra Teresa Cignarella, Teresa Lynn, Ines Rehbein, Josef Ruppenhofer, Djamé Seddah, Amir Zeldes
| Challenge: | Despite the increasing number of contributions on Part-of-Speech tagging and parsing, automatic processing of user-generated content (UGC) still represents a challenging task. |
| Approach: | They propose a set of guidelines for the annotation of user-generated texts within the Universal Dependencies framework. |
| Outcome: | The proposed annotation guidelines promote cross-linguistic consistency, which has always been in the spirit of UD. |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
What Meaning-Form Correlation Has to Compose With: A Study of MFC on Artificial and Natural Language (2020.coling-main)
Copied to clipboard
| Challenge: | Compositionality is a widely discussed property of natural languages, although its exact definition has been elusive. |
| Approach: | They propose that compositionality can be measured by measuring meaning-form correlation . they analyze three sets of languages: artificial toy languages tailored to be compositional . |
| Outcome: | The proposed method can assess compositionality on three sets of languages . linguistic phenomena such as synonymy and ungrounded stop-words weigh on the results . |
Cost-Effective Discourse Annotation in the Prague Czech–English Dependency Treebank (2024.lrec-main)
Copied to clipboard
| Challenge: | a method for obtaining a high-quality annotation of explicit discourse relations is a resource-demanding task. |
| Approach: | They propose a method for obtaining a high-quality annotation of explicit discourse relations in the Czech part of the Prague Czech–English Dependency Treebank. |
| Outcome: | The proposed method solves the problem of identifying discrepancies between the annotations in the Czech part of the Penn Treebank. |