Papers by Daniel Zeman
Towards a Unified Taxonomy of Deep Syntactic Relations (2024.lrec-main)
Copied to clipboard
| Challenge: | Currently, UD is the standard for morphology and surface syntax annotations, but it is only one step towards natural language understanding. |
| Approach: | They propose to use a set of universal semantic role labels for morphology and surface syntax in four Indo-European and one Uralic languages to analyze the data. |
| Outcome: | The proposed set of universal semantic role labels is based on the data from four Indo-European and one Uralic languages. |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
Do UD Trees Match Mention Spans in Coreference Annotations? (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to annotate mention spans are based on delimiting token intervals, but there is no syntactic representation of the mention span. |
| Approach: | They propose to integrate coreference annotation with syntactic annotation to make them convergent in the long term. |
| Outcome: | The proposed approach could be advantageous in the long term, the authors argue. |
Universal Anaphora: The First Three Years (2024.lrec-main)
Copied to clipboard
Massimo Poesio, Maciej Ogrodniczuk, Vincent Ng, Sameer Pradhan, Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Amir Zeldes, Anna Nedoluzhko, Michal Novák, Martin Popel, Zdeněk Žabokrtský, Daniel Zeman
| Challenge: | Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by expanding the aspects of anaphonic interpretation which are or can be reliably annotated in an anagraphic corpora. |
| Approach: | They propose to develop a standard for anaphoric annotations and a method for evaluating models that can carry out this type of interpretation. |
| Outcome: | The Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by producing unified standards to annotate and encode annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation. |
Parse Me if You Can: Artificial Treebanks for Parsing Experiments on Elliptical Constructions (L18-1)
Copied to clipboard
| Challenge: | ellipsis is a phenomenon present in many natural languages, but it complicates syntactic parsing of the content that is not omitted. |
| Approach: | They analyze outputs of state-of-the-art parsers to learn about parsing accuracy and typical errors from the perspective of elliptical constructions. |
| Outcome: | The proposed treebank is a semi-artificially constructed treebank of ellipsis. |
CorefUD 1.0: Coreference Meets Universal Dependencies (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in standardization for annotated language resources have led to successful large scale efforts, such as the Universal Dependencies (UD) project for multilingual syntactically annotized data. |
| Approach: | They propose a multilingual collection of corpora and a standardized format for coreference resolution compatible with morphosyntactic annotations in the UD framework. |
| Outcome: | The proposed framework is compatible with morphosyntactic annotations and includes facilities for related tasks such as named entity recognition. |
Yorùbá Dependency Treebank (YTB) (2020.lrec-1)
Copied to clipboard
| Challenge: | Low-resource languages present enormous NLP opportunities as well as varying degrees of difficulties. |
| Approach: | They propose to use the Yoruba Bible treebank to apply a new grammar formalism to the language by examining the use of universal dependency annotations. |
| Outcome: | The treebank of hand-annotated parts of the Yoruba Bible provides an avenue for dependency analysis of the language; the application of a new grammar formalism to the language. |