Papers by Alena Fenogenova
Humans Keep It One Hundred: an Overview of AI Journey (2020.lrec-1)
Copied to clipboard
Tatiana Shavrina, Anton Emelyanov, Alena Fenogenova, Vadim Fomin, Vladislav Mikhailov, Andrey Evlampiev, Valentin Malykh, Vladimir Larin, Alex Natekin, Aleksandr Vatulin, Peter Romov, Daniil Anastasiev, Nikolai Zinov, Andrey Chertok
| Challenge: | Artificial General Intelligence (AGI) is showing growing performance in numerous applications - beating human performance in Chess and Go, using knowledge bases and text sources to answer questions and even pass human examination. |
| Approach: | They propose to use knowledge bases and text sources to answer questions to improve AI performance on knowledge bases, reasoning and text generation. |
| Outcome: | The proposed AI Journey system passed the final native language exam in Russian with a high score of 69%, with 68% being an average human result. |
A Family of Pretrained Transformer Language Models for Russian (2024.lrec-main)
Copied to clipboard
Dmitry Zmitrovich, Aleksandr Abramov, Andrey Kalmykov, Vitaly Kadulin, Maria Tikhonova, Ekaterina Taktasheva, Danil Astafurov, Mark Baushenko, Artem Snegirev, Tatiana Shavrina, Sergei S. Markov, Vladislav Mikhailov, Alena Fenogenova
| Challenge: | Developing Transformer language models for the Russian language has received little attention . most of these LMs are developed for English, which imposes substantial constraints on the potential of the language technologies. |
| Approach: | They propose to release 13 Russian Transformer language models that span three languages . they aim to broaden the scope of NLP research directions and develop industrial solutions for the Russian language. |
| Outcome: | The proposed models are based on Russian language datasets and benchmarks. |
DRAGOn: Designing RAG On Periodically Updated Corpus (2026.eacl-srw)
Copied to clipboard
Fedor Chernogorskii, Sergei Averkiev, Liliya Kudraleeva, Zaven Martirosian, Maria Tikhonova, Valentin Malykh, Alena Fenogenova
| Challenge: | Existing methods for evaluating RAG systems are labor-intensive and difficult to maintain. |
| Approach: | They propose a method to design a RAG benchmark on a regularly updated corpus. |
| Outcome: | The proposed method uses a regularly updated corpus to evaluate RAG models. |
A Methodology for Generative Spelling Correction via Natural Spelling Errors Emulation across Multiple Domains and Languages (2024.findings-eacl)
Copied to clipboard
Nikita Martynov, Mark Baushenko, Anastasia Kozlova, Katerina Kolomeytseva, Aleksandr Abramov, Alena Fenogenova
| Challenge: | Recent advances in large language models have shown impressive text generation and language understanding capabilities, evident in benchmarks like SuperGLUE, GEM, BigBench etc. |
| Approach: | They propose a method for generative spelling correction that can be extended to any language with minor changes. |
| Outcome: | The proposed method can be extended to any language with minor changes, and is based on a set of generative models with a single-domain and multi-domain test sets. |
Read and Reason with MuSeRC and RuCoS: Datasets for Machine Reading Comprehension for Russian (2020.coling-main)
Copied to clipboard
| Challenge: | MRC in other languages, including Russian, has not been well-addressed due to the lack of high-quality and large-scale datasets. |
| Approach: | They propose two Russian machine reading comprehension datasets that require reasoning over multiple sentences and commonsense knowledge to infer the answer. |
| Outcome: | The proposed datasets are more complex than the original ones for Russian . the results show that the proposed models are challenging for advanced models . |
FiMMIA: scaling semantic perturbation-based membership inference across modalities (2026.eacl-demo)
Copied to clipboard
| Challenge: | Membership Inference attacks aim to determine whether a specific data point was included in the training set of a target model. |
| Approach: | They propose to train a neural network to analyze the target model’s behavior on perturbed inputs, capturing interactions between semantic domains and loss values on members and non-members in the local neighborhood of each sample. |
| Outcome: | The proposed methods can detect distribution shifts in existing datasets and release a baseline pipeline to detect them. |
TAPE: Assessing Few-shot Russian Language Understanding (2022.findings-emnlp)
Copied to clipboard
Ekaterina Taktasheva, Tatiana Shavrina, Alena Fenogenova, Denis Shevelev, Nadezhda Katricheva, Maria Tikhonova, Albina Akhmetgareeva, Oleg Zinkevich, Anastasiia Bashmakova, Svetlana Iordanskaia, Alena Spiridonova, Valentina Kurenshchikova, Ekaterina Artemova, Vladislav Mikhailov
| Challenge: | Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes, but lacks standardized evaluation suites for non-English languages. |
| Approach: | They propose a novel benchmark that includes six more complex NLU tasks for Russian, covering multi-hop reasoning, ethical concepts, logic and commonsense knowledge. |
| Outcome: | The proposed benchmark includes six more complex NLU tasks for Russian, covering multi-hop reasoning, ethical concepts, logic and commonsense knowledge. |
GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture (2025.acl-demo)
Copied to clipboard
Valentin Mamedov, Evgenii Kosarev, Gregory Leleytner, Ilya Shchuckin, Valeriy Berezovskiy, Daniil Smirnov, Dmitry Kozlov, Sergei Averkiev, Lukyanenko Ivan, Aleksandr Proshunin, Ainur Israfilova, Ivan Baskov, Artem Chervyakov, Emil Shakirov, Mikhail Kolesov, Daria Khomich, Daria Latortseva, Sergei Porkhun, Yury Fedorov, Oleg Kutuzov, Polina Kudriavtseva, Sofiia Soldatova, Kolodin Egor, Stanislav Pyatkin, Dzmitry Menshykh, Grafov Sergei IUrevich, Eldar Damirov, Vladimir Karlov, Ruslan Gaitukiev, Arkadiy Shatenov, Alena Fenogenova, Nikita Savushkin, Fedor Minkin
| Challenge: | generative large language models have become crucial for modern NLP research and applications across multiple languages. |
| Approach: | They introduce the GigaChat family of Russian LLMs, available in various sizes . they evaluate their performance on Russian and English benchmarks and compare them with multilingual analogs . |
| Outcome: | The proposed model family is available in various sizes and is tested on Russian and English benchmarks. |
2Columns1Row: A Russian Benchmark for Textual and Multimodal Table Understanding and Reasoning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | 2Columns1Row is the first open-source benchmark for the table understanding task in Russian. |
| Approach: | They propose a benchmark for table understanding in Russian using textual and multimodal inputs. |
| Outcome: | The proposed benchmark evaluates models' ability to reason about relationships between rows and columns in tables using text-only and multimodal approaches. |
MERA: A Comprehensive LLM Evaluation in Russian (2024.acl-long)
Copied to clipboard
Alena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova, Maria Tikhonova, Albina Akhmetgareeva, Anton Emelyanov, Denis Shevelev, Pavel Lebedev, Leonid Sinev, Ulyana Isaeva, Katerina Kolomeytseva, Daniil Moskovskiy, Elizaveta Goncharova, Nikita Savushkin, Polina Mikhailova, Anastasia Minaeva, Denis Dimitrov, Alexander Panchenko, Sergey Markov
| Challenge: | Recent advances in foundation models have led to the emergence of powerful Large Language Models (LLMs), which showcase unprecedented tasksolving capabilities. |
| Approach: | They propose a method to evaluate FMs and LMs in fixed zero- and few-shot instruction settings that can be extended to other modalities. |
| Outcome: | The proposed evaluation methodology includes an open-source code base and a leaderboard with a submission system. |
The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design (2025.naacl-long)
Copied to clipboard
| Challenge: | Embedding models are used in tasks such as information retrieval and semantic textual similarity. |
| Approach: | They propose a new Russian-focused embedding model called ru-en-RoSBERTa and a benchmark for Russian language . they propose to use the roMTEB benchmark to assess Russian and multilingual models . |
| Outcome: | The proposed model achieves results that are on par with state-of-the-art models in Russian. |
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs (2024.emnlp-main)
Copied to clipboard
Ekaterina Taktasheva, Maxim Bazhukov, Kirill Koncha, Alena Fenogenova, Ekaterina Artemova, Vladislav Mikhailov
| Challenge: | Existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. |
| Approach: | They propose to use a Russian benchmark of linguistic minimal pairs to evaluate grammatical knowledge of language models. |
| Outcome: | The proposed benchmark includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon. |
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks (2025.emnlp-demos)
Copied to clipboard
Adamenko Pavel, Ivanov Mikhail, Aidar Valeev, Rodion Levichev, Pavel Zadorozhny, Ivan Lopatin, Dmitrii Babaev, Alena Fenogenova, Valentin Malykh
| Challenge: | SWE-bench is a static benchmark that collects only once and never updates. |
| Approach: | They propose a dynamic, continuously updated benchmark to address data contamination issues by collecting real-world GitHub issues and rigorous quality validation. |
| Outcome: | The proposed benchmarks are based on a dataset of 2,294 GitHub issues and their corresponding pull requests (PRs) the static nature of the benchmarks makes it hard to distinguish meaningful progress. |
Multimodal Evaluation of Russian-language Architectures (2026.eacl-long)
Copied to clipboard
Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov, Artem Safin, Maria Tikhonova, Alexander Kharitonov, Yulia Lyakh, Petr Surovtsev, Denis Shevelev, Vildan Saburov, Vasily Konovalov, Elisei Rykov, Ivan Sviridov, Amina Miftakhova, Ilseyar Alimova, Alexander Panchenko, Alexander Kapitanov, Alena Fenogenova
| Challenge: | Multimodal large language models (MLLMs) are at the center of research attention, yet intelligence, limitations, and risks remain insufficiently understood. |
| Approach: | They propose an open multimodal evaluation framework for Russian-spoken architectures . the framework is instruction-based and includes 18 newly constructed evaluation tasks . |
| Outcome: | The proposed framework provides a replicable methodology for constructing multimodal benchmarks in Russian-spoken architectures. |
RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark (2020.emnlp-main)
Copied to clipboard
Tatiana Shavrina, Alena Fenogenova, Emelyanov Anton, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, Andrey Evlampiev
| Challenge: | Modern scientific methodology is beginning to explore universal transformers as an independent object of study. |
| Approach: | They propose a Russian general language understanding evaluation benchmark - Russian SuperGLUE . they provide a benchmark of nine tasks, human level evaluation and a leaderboard for the Russian language . |
| Outcome: | The proposed benchmark provides nine tasks for the Russian language and human level evaluation and leaderboard of transformer models. |