Papers by Zdeněk Žabokrtský

8 papers
Constructing a Lexical Resource of Russian Derivational Morphology (2022.lrec-1)

Copied to clipboard

Challenge: In Natural Language Processing of Russian, the inflection is satisfactorily processed, but there are only a few machine-trackable resources that capture derivations .
Approach: They propose to use machine-learning methods to improve Russian inflection and derivational resources by using a database of more than 300 thousand lexemes and 164 thousand binary derivations.
Outcome: The proposed method includes more than 300 thousand lexemes connected with more than 164 thousand binary derivational relations.
Using Adversarial Examples in Natural Language Processing (L18-1)

Copied to clipboard

Challenge: Recent advances in machine learning have led to the use of adversarial examples in training of neural networks.
Approach: They investigate the effect of using adversarial examples during training of recurrent neural networks whose text input is in the form of a sequence of word/character embeddings.
Outcome: The proposed method provides regularization effect and enables training of models with greater number of parameters without overfitting.
Towards Universal Segmentations: UniSegments 1.0 (2022.lrec-1)

Copied to clipboard

Challenge: Existing data resources for morphological segmentation are limited to 32 languages . a large number of word forms exist, with some sub-parts being "recycled" many times .
Approach: They propose a multilingual data resource for morphological segmentation in 32 languages . they analyze diversity of how individual linguistic phenomena are captured across them .
Outcome: The proposed scheme is based on 17 existing data resources relevant for segmentation in 32 languages.
Do UD Trees Match Mention Spans in Coreference Annotations? (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to annotate mention spans are based on delimiting token intervals, but there is no syntactic representation of the mention span.
Approach: They propose to integrate coreference annotation with syntactic annotation to make them convergent in the long term.
Outcome: The proposed approach could be advantageous in the long term, the authors argue.
Universal Anaphora: The First Three Years (2024.lrec-main)

Copied to clipboard

Challenge: Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by expanding the aspects of anaphonic interpretation which are or can be reliably annotated in an anagraphic corpora.
Approach: They propose to develop a standard for anaphoric annotations and a method for evaluating models that can carry out this type of interpretation.
Outcome: The Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by producing unified standards to annotate and encode annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation.
CorefUD 1.0: Coreference Meets Universal Dependencies (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in standardization for annotated language resources have led to successful large scale efforts, such as the Universal Dependencies (UD) project for multilingual syntactically annotized data.
Approach: They propose a multilingual collection of corpora and a standardized format for coreference resolution compatible with morphosyntactic annotations in the UD framework.
Outcome: The proposed framework is compatible with morphosyntactic annotations and includes facilities for related tasks such as named entity recognition.
SumeCzech: Large Czech News-Based Summarization Dataset (L18-1)

Copied to clipboard

Challenge: Summarization of documents is a well-studied NLP task, but only a few datasets are available for Czech.
Approach: They propose to use a Czech news-based summarization dataset to evaluate document summarizing . they propose a language-agnostic variant of the ROUGE metric to enable automatic evaluation .
Outcome: The proposed dataset contains more than a million Czech news articles . the proposed approach is strong abstractive and language-agnostic .
Semi-Automatic Construction of Word-Formation Networks (for Polish and Spanish) (L18-1)

Copied to clipboard

Challenge: a semi-automatic method for the construction of derivational networks is proposed . the proposed method is general enough to be adopted for other languages .
Approach: They propose a semi-automatic method for the construction of derivational networks using a sequential pattern mining technique.
Outcome: The proposed method is general enough to be adopted for other languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations