Papers by Andrey Sakhovskiy
Biomedical Entity Representation with Graph-Augmented Multi-Objective Transformer (2024.findings-naacl)
Copied to clipboard
| Challenge: | Modern biomedical concept representations are mostly trained on synonymous concept names from a biomedically knowledge base graph, ignoring the inter-concept interactions and a concept’s local neighborhood. |
| Approach: | They propose a Graph-Augmented Multi-Objective Transformer which captures both inter-concept and intra-conception interactions from the multilingual UMLS graph. |
| Outcome: | The proposed model captures inter- and intra-concept interactions from the multilingual UMLS graph using pre-trained language models and graph neural networks. |
Biomedical Concept Normalization over Nested Entities with Partial UMLS Terminology in Russian (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing annotations in Russian do not include all entities, but only a small fraction of them are labeled in English. |
| Approach: | They present a manually annotated PubMed abstract dataset for concept normalization in Russian. |
| Outcome: | The proposed model improves on nested named entities in a zero-shot setting on bilingual terminology. |
InkSight: Towards AI-Aided Historical Manuscript Analysis (2026.eacl-demo)
Copied to clipboard
Andrey Sakhovskiy, Ivan Ulitin, Emilia Bojarskaja, Vladimir Kokh, Ruslan Murtazin, Maxim Novopoltsev, Semen Budennyy
| Challenge: | Large-scale scientific research on medieval Arabic manuscripts remains challenging due to the need for advanced paleographic and linguistic training and the lack of assisting software. |
| Approach: | They propose an end-to-end Arabic manuscript analysis tool for manuscript-based analytics and research hypothesis testing. |
| Outcome: | The proposed tool overcomes the limitations of existing tools and can be used in large-scale scientific research. |
RuCCoD: Towards Automated ICD Coding in Russian (2025.emnlp-main)
Copied to clipboard
Alexandr Nesterov, Andrey Sakhovskiy, Ivan Sviridov, Airat Valiev, Vladimir Makharev, Petr Anokhin, Galina Zubkova, Elena Tutubalina
| Challenge: | a new dataset for clinical coding in Russian is available for download . human coders must navigate a wide array of medical terminology and time pressures . |
| Approach: | They present a new dataset for ICD coding in Russian, a language with limited biomedical resources. |
| Outcome: | The proposed model improves accuracy on an in-house EHR dataset from 2017 to 2021. |
Lost in Translation: Chemical Language Models and the Misunderstanding of Molecule Structures (2024.findings-emnlp)
Copied to clipboard
Veronika Ganeeva, Andrey Sakhovskiy, Kuzma Khrabrov, Andrey Savchenko, Artur Kadurin, Elena Tutubalina
| Challenge: | chemistry and natural language processing (NLP) have advanced drug discovery. |
| Approach: | They propose a framework for assessment of Chemistry LMs of different natures that relies on augmentations that preserve an underlying chemical. |
| Outcome: | The proposed framework relies on augmentations that preserve an underlying chemical, such as kekulization and cycle replacements. |