Papers by Artem Snegirev

2 papers
A Family of Pretrained Transformer Language Models for Russian (2024.lrec-main)

Copied to clipboard

Challenge: Developing Transformer language models for the Russian language has received little attention . most of these LMs are developed for English, which imposes substantial constraints on the potential of the language technologies.
Approach: They propose to release 13 Russian Transformer language models that span three languages . they aim to broaden the scope of NLP research directions and develop industrial solutions for the Russian language.
Outcome: The proposed models are based on Russian language datasets and benchmarks.
The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design (2025.naacl-long)

Copied to clipboard

Challenge: Embedding models are used in tasks such as information retrieval and semantic textual similarity.
Approach: They propose a new Russian-focused embedding model called ru-en-RoSBERTa and a benchmark for Russian language . they propose to use the roMTEB benchmark to assess Russian and multilingual models .
Outcome: The proposed model achieves results that are on par with state-of-the-art models in Russian.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations