Challenge: DerivBase.Ru is a high-coverage derivational morphology resource for Russian language that can be used for many tasks such as paraphrases and plagiarism detection.
Approach: They propose a rule-based framework for deriving Russian words using a derivational morphology resource called DerivBase.Ru.
Outcome: The proposed framework can be used to derivate words from a dictionary in Russian and German.

Similar Papers

Constructing a Lexical Resource of Russian Derivational Morphology (2022.lrec-1)

Copied to clipboard

Challenge: In Natural Language Processing of Russian, the inflection is satisfactorily processed, but there are only a few machine-trackable resources that capture derivations .
Approach: They propose to use machine-learning methods to improve Russian inflection and derivational resources by using a database of more than 300 thousand lexemes and 164 thousand binary derivations.
Outcome: The proposed method includes more than 300 thousand lexemes connected with more than 164 thousand binary derivational relations.
Semi-Automatic Construction of Word-Formation Networks (for Polish and Spanish) (L18-1)

Copied to clipboard

Challenge: a semi-automatic method for the construction of derivational networks is proposed . the proposed method is general enough to be adopted for other languages .
Approach: They propose a semi-automatic method for the construction of derivational networks using a sequential pattern mining technique.
Outcome: The proposed method is general enough to be adopted for other languages.
Computational Etymology and Word Emergence (2020.lrec-1)

Copied to clipboard

Challenge: etymology is the study of words' origins.
Approach: They develop an extensible Wiktionary parser that predicts the etymology of a word across the full range of ethymological types and languages in Wiktionaries.
Outcome: The proposed parser predicts the etymology of a word across the full range of ethymologies and languages in Wiktionary, and shows the application of tymatics in modeling this phenomenon.
Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages (2022.coling-1)

Copied to clipboard

Challenge: a detailed empirical case study of out-of-vocabulary words in modern text is presented . unfamiliar words cause trouble for machine processing or comprehension of text, authors say .
Approach: They propose a detailed empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian . they apply a multi-faceted distributional analysis of the underlying word-formation processes to characterize the residual vocabulary .
Outcome: The proposed method can be used to aid in compositional translation, parsing, language modeling, and other NLP tasks.
OOVs in the Spotlight: How to Inflect Them? (2024.lrec-main)

Copied to clipboard

Challenge: Inflection is a process of word formation in which a base word form (lemma) is modified to express grammatical categories.
Approach: They develop a retrograde model and two sequence-to-sequence models based on LSTM and Transformer.
Outcome: The proposed systems outperform the existing systems on 9 out of 16 languages in the OOV evaluation.
A Graph Auto-encoder Model of Derivational Morphology (2020.acl-main)

Copied to clipboard

Challenge: Existing words that conform to morphological patterns of a language differ in how likely they are to be actually created by speakers.
Approach: They propose to model the morphological well-formedness of derivatives by combining syntactic and semantic information with associative information from the mental lexicon.
Outcome: The proposed model models the morphological well-formedness of derivatives in English .
BERT-like Models for Slavic Morpheme Segmentation (2025.acl-long)

Copied to clipboard

Challenge: Existing morpheme segmentation algorithms for Slavic languages have been improved but performance is still low for words with roots not present in training data.
Approach: They propose to fine-tune BERT-like models for morpheme segmentation using data from Belarusian, Czech, and Russian to account for word semantics.
Outcome: The proposed models outperform all previous approaches in Czech and Russian, with word-level accuracy of 92.5-95.1%.
New Dataset and Strong Baselines for the Grammatical Error Correction of Russian (2021.findings-acl)

Copied to clipboard

Challenge: a new resource is created to evaluate grammatical error correction models in English . a subset of the dataset is annotated in Russian, which is hard to come by and expensive to annotate .
Approach: They develop an annotated learner corpus of Russian extracted from the Lang-8 website.
Outcome: The proposed dataset is compared against two state-of-the-art grammatical error correction models . the results show that the created corpus is more diverse than the existing one .
Transactions of the Association for Computational Linguistics, Volume 8 (2020.tacl-1)

Copied to clipboard

Challenge: null
Approach: null
Outcome: null
A Distributional and Orthographic Aggregation Model for English Derivational Morphology (P18-1)

Copied to clipboard

Challenge: Existing approaches to derived word generation model derivational morphology to generate words with particular semantics are not effective.
Approach: They propose a novel aggregation model that learns derivational transformations as orthographic functions and as functions in distributional word embedding space.
Outcome: The proposed model learns to choose between the hypothesis of each system and the hypothesis from the model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations