Papers by Artem Snegirev
A Family of Pretrained Transformer Language Models for Russian (2024.lrec-main)
Copied to clipboard
Dmitry Zmitrovich, Aleksandr Abramov, Andrey Kalmykov, Vitaly Kadulin, Maria Tikhonova, Ekaterina Taktasheva, Danil Astafurov, Mark Baushenko, Artem Snegirev, Tatiana Shavrina, Sergei S. Markov, Vladislav Mikhailov, Alena Fenogenova
| Challenge: | Developing Transformer language models for the Russian language has received little attention . most of these LMs are developed for English, which imposes substantial constraints on the potential of the language technologies. |
| Approach: | They propose to release 13 Russian Transformer language models that span three languages . they aim to broaden the scope of NLP research directions and develop industrial solutions for the Russian language. |
| Outcome: | The proposed models are based on Russian language datasets and benchmarks. |
The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design (2025.naacl-long)
Copied to clipboard
| Challenge: | Embedding models are used in tasks such as information retrieval and semantic textual similarity. |
| Approach: | They propose a new Russian-focused embedding model called ru-en-RoSBERTa and a benchmark for Russian language . they propose to use the roMTEB benchmark to assess Russian and multilingual models . |
| Outcome: | The proposed model achieves results that are on par with state-of-the-art models in Russian. |