Papers by Mark Baushenko
A Family of Pretrained Transformer Language Models for Russian (2024.lrec-main)
Copied to clipboard
Dmitry Zmitrovich, Aleksandr Abramov, Andrey Kalmykov, Vitaly Kadulin, Maria Tikhonova, Ekaterina Taktasheva, Danil Astafurov, Mark Baushenko, Artem Snegirev, Tatiana Shavrina, Sergei S. Markov, Vladislav Mikhailov, Alena Fenogenova
| Challenge: | Developing Transformer language models for the Russian language has received little attention . most of these LMs are developed for English, which imposes substantial constraints on the potential of the language technologies. |
| Approach: | They propose to release 13 Russian Transformer language models that span three languages . they aim to broaden the scope of NLP research directions and develop industrial solutions for the Russian language. |
| Outcome: | The proposed models are based on Russian language datasets and benchmarks. |
A Methodology for Generative Spelling Correction via Natural Spelling Errors Emulation across Multiple Domains and Languages (2024.findings-eacl)
Copied to clipboard
Nikita Martynov, Mark Baushenko, Anastasia Kozlova, Katerina Kolomeytseva, Aleksandr Abramov, Alena Fenogenova
| Challenge: | Recent advances in large language models have shown impressive text generation and language understanding capabilities, evident in benchmarks like SuperGLUE, GEM, BigBench etc. |
| Approach: | They propose a method for generative spelling correction that can be extended to any language with minor changes. |
| Outcome: | The proposed method can be extended to any language with minor changes, and is based on a set of generative models with a single-domain and multi-domain test sets. |