Papers by Denis Shevelev
Read and Reason with MuSeRC and RuCoS: Datasets for Machine Reading Comprehension for Russian (2020.coling-main)
Copied to clipboard
| Challenge: | MRC in other languages, including Russian, has not been well-addressed due to the lack of high-quality and large-scale datasets. |
| Approach: | They propose two Russian machine reading comprehension datasets that require reasoning over multiple sentences and commonsense knowledge to infer the answer. |
| Outcome: | The proposed datasets are more complex than the original ones for Russian . the results show that the proposed models are challenging for advanced models . |
TAPE: Assessing Few-shot Russian Language Understanding (2022.findings-emnlp)
Copied to clipboard
Ekaterina Taktasheva, Tatiana Shavrina, Alena Fenogenova, Denis Shevelev, Nadezhda Katricheva, Maria Tikhonova, Albina Akhmetgareeva, Oleg Zinkevich, Anastasiia Bashmakova, Svetlana Iordanskaia, Alena Spiridonova, Valentina Kurenshchikova, Ekaterina Artemova, Vladislav Mikhailov
| Challenge: | Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes, but lacks standardized evaluation suites for non-English languages. |
| Approach: | They propose a novel benchmark that includes six more complex NLU tasks for Russian, covering multi-hop reasoning, ethical concepts, logic and commonsense knowledge. |
| Outcome: | The proposed benchmark includes six more complex NLU tasks for Russian, covering multi-hop reasoning, ethical concepts, logic and commonsense knowledge. |
MERA: A Comprehensive LLM Evaluation in Russian (2024.acl-long)
Copied to clipboard
Alena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova, Maria Tikhonova, Albina Akhmetgareeva, Anton Emelyanov, Denis Shevelev, Pavel Lebedev, Leonid Sinev, Ulyana Isaeva, Katerina Kolomeytseva, Daniil Moskovskiy, Elizaveta Goncharova, Nikita Savushkin, Polina Mikhailova, Anastasia Minaeva, Denis Dimitrov, Alexander Panchenko, Sergey Markov
| Challenge: | Recent advances in foundation models have led to the emergence of powerful Large Language Models (LLMs), which showcase unprecedented tasksolving capabilities. |
| Approach: | They propose a method to evaluate FMs and LMs in fixed zero- and few-shot instruction settings that can be extended to other modalities. |
| Outcome: | The proposed evaluation methodology includes an open-source code base and a leaderboard with a submission system. |
Multimodal Evaluation of Russian-language Architectures (2026.eacl-long)
Copied to clipboard
Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov, Artem Safin, Maria Tikhonova, Alexander Kharitonov, Yulia Lyakh, Petr Surovtsev, Denis Shevelev, Vildan Saburov, Vasily Konovalov, Elisei Rykov, Ivan Sviridov, Amina Miftakhova, Ilseyar Alimova, Alexander Panchenko, Alexander Kapitanov, Alena Fenogenova
| Challenge: | Multimodal large language models (MLLMs) are at the center of research attention, yet intelligence, limitations, and risks remain insufficiently understood. |
| Approach: | They propose an open multimodal evaluation framework for Russian-spoken architectures . the framework is instruction-based and includes 18 newly constructed evaluation tasks . |
| Outcome: | The proposed framework provides a replicable methodology for constructing multimodal benchmarks in Russian-spoken architectures. |
RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark (2020.emnlp-main)
Copied to clipboard
Tatiana Shavrina, Alena Fenogenova, Emelyanov Anton, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, Andrey Evlampiev
| Challenge: | Modern scientific methodology is beginning to explore universal transformers as an independent object of study. |
| Approach: | They propose a Russian general language understanding evaluation benchmark - Russian SuperGLUE . they provide a benchmark of nine tasks, human level evaluation and a leaderboard for the Russian language . |
| Outcome: | The proposed benchmark provides nine tasks for the Russian language and human level evaluation and leaderboard of transformer models. |