| Challenge: | DerivBase.Ru is a high-coverage derivational morphology resource for Russian language that can be used for many tasks such as paraphrases and plagiarism detection. |
| Approach: | They propose a rule-based framework for deriving Russian words using a derivational morphology resource called DerivBase.Ru. |
| Outcome: | The proposed framework can be used to derivate words from a dictionary in Russian and German. |
Similar Papers
Constructing a Lexical Resource of Russian Derivational Morphology (2022.lrec-1)
Copied to clipboard
| Challenge: | In Natural Language Processing of Russian, the inflection is satisfactorily processed, but there are only a few machine-trackable resources that capture derivations . |
| Approach: | They propose to use machine-learning methods to improve Russian inflection and derivational resources by using a database of more than 300 thousand lexemes and 164 thousand binary derivations. |
| Outcome: | The proposed method includes more than 300 thousand lexemes connected with more than 164 thousand binary derivational relations. |
Semi-Automatic Construction of Word-Formation Networks (for Polish and Spanish) (L18-1)
Copied to clipboard
| Challenge: | a semi-automatic method for the construction of derivational networks is proposed . the proposed method is general enough to be adopted for other languages . |
| Approach: | They propose a semi-automatic method for the construction of derivational networks using a sequential pattern mining technique. |
| Outcome: | The proposed method is general enough to be adopted for other languages. |
Computational Etymology and Word Emergence (2020.lrec-1)
Copied to clipboard
| Challenge: | etymology is the study of words' origins. |
| Approach: | They develop an extensible Wiktionary parser that predicts the etymology of a word across the full range of ethymological types and languages in Wiktionaries. |
| Outcome: | The proposed parser predicts the etymology of a word across the full range of ethymologies and languages in Wiktionary, and shows the application of tymatics in modeling this phenomenon. |
Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages (2022.coling-1)
Copied to clipboard
| Challenge: | a detailed empirical case study of out-of-vocabulary words in modern text is presented . unfamiliar words cause trouble for machine processing or comprehension of text, authors say . |
| Approach: | They propose a detailed empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian . they apply a multi-faceted distributional analysis of the underlying word-formation processes to characterize the residual vocabulary . |
| Outcome: | The proposed method can be used to aid in compositional translation, parsing, language modeling, and other NLP tasks. |
OOVs in the Spotlight: How to Inflect Them? (2024.lrec-main)
Copied to clipboard
| Challenge: | Inflection is a process of word formation in which a base word form (lemma) is modified to express grammatical categories. |
| Approach: | They develop a retrograde model and two sequence-to-sequence models based on LSTM and Transformer. |
| Outcome: | The proposed systems outperform the existing systems on 9 out of 16 languages in the OOV evaluation. |
A Graph Auto-encoder Model of Derivational Morphology (2020.acl-main)
Copied to clipboard
| Challenge: | Existing words that conform to morphological patterns of a language differ in how likely they are to be actually created by speakers. |
| Approach: | They propose to model the morphological well-formedness of derivatives by combining syntactic and semantic information with associative information from the mental lexicon. |
| Outcome: | The proposed model models the morphological well-formedness of derivatives in English . |
BERT-like Models for Slavic Morpheme Segmentation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing morpheme segmentation algorithms for Slavic languages have been improved but performance is still low for words with roots not present in training data. |
| Approach: | They propose to fine-tune BERT-like models for morpheme segmentation using data from Belarusian, Czech, and Russian to account for word semantics. |
| Outcome: | The proposed models outperform all previous approaches in Czech and Russian, with word-level accuracy of 92.5-95.1%. |
New Dataset and Strong Baselines for the Grammatical Error Correction of Russian (2021.findings-acl)
Copied to clipboard
| Challenge: | a new resource is created to evaluate grammatical error correction models in English . a subset of the dataset is annotated in Russian, which is hard to come by and expensive to annotate . |
| Approach: | They develop an annotated learner corpus of Russian extracted from the Lang-8 website. |
| Outcome: | The proposed dataset is compared against two state-of-the-art grammatical error correction models . the results show that the created corpus is more diverse than the existing one . |
Transactions of the Association for Computational Linguistics, Volume 8 (2020.tacl-1)
Copied to clipboard
| Challenge: | null |
| Approach: | null |
| Outcome: | null |
A Distributional and Orthographic Aggregation Model for English Derivational Morphology (P18-1)
Copied to clipboard
| Challenge: | Existing approaches to derived word generation model derivational morphology to generate words with particular semantics are not effective. |
| Approach: | They propose a novel aggregation model that learns derivational transformations as orthographic functions and as functions in distributional word embedding space. |
| Outcome: | The proposed model learns to choose between the hypothesis of each system and the hypothesis from the model. |