Papers by Andrey Sakhovskiy

5 papers
Biomedical Entity Representation with Graph-Augmented Multi-Objective Transformer (2024.findings-naacl)

Copied to clipboard

Challenge: Modern biomedical concept representations are mostly trained on synonymous concept names from a biomedically knowledge base graph, ignoring the inter-concept interactions and a concept’s local neighborhood.
Approach: They propose a Graph-Augmented Multi-Objective Transformer which captures both inter-concept and intra-conception interactions from the multilingual UMLS graph.
Outcome: The proposed model captures inter- and intra-concept interactions from the multilingual UMLS graph using pre-trained language models and graph neural networks.
Biomedical Concept Normalization over Nested Entities with Partial UMLS Terminology in Russian (2024.lrec-main)

Copied to clipboard

Challenge: Existing annotations in Russian do not include all entities, but only a small fraction of them are labeled in English.
Approach: They present a manually annotated PubMed abstract dataset for concept normalization in Russian.
Outcome: The proposed model improves on nested named entities in a zero-shot setting on bilingual terminology.
InkSight: Towards AI-Aided Historical Manuscript Analysis (2026.eacl-demo)

Copied to clipboard

Challenge: Large-scale scientific research on medieval Arabic manuscripts remains challenging due to the need for advanced paleographic and linguistic training and the lack of assisting software.
Approach: They propose an end-to-end Arabic manuscript analysis tool for manuscript-based analytics and research hypothesis testing.
Outcome: The proposed tool overcomes the limitations of existing tools and can be used in large-scale scientific research.
RuCCoD: Towards Automated ICD Coding in Russian (2025.emnlp-main)

Copied to clipboard

Challenge: a new dataset for clinical coding in Russian is available for download . human coders must navigate a wide array of medical terminology and time pressures .
Approach: They present a new dataset for ICD coding in Russian, a language with limited biomedical resources.
Outcome: The proposed model improves accuracy on an in-house EHR dataset from 2017 to 2021.
Lost in Translation: Chemical Language Models and the Misunderstanding of Molecule Structures (2024.findings-emnlp)

Copied to clipboard

Challenge: chemistry and natural language processing (NLP) have advanced drug discovery.
Approach: They propose a framework for assessment of Chemistry LMs of different natures that relies on augmentations that preserve an underlying chemical.
Outcome: The proposed framework relies on augmentations that preserve an underlying chemical, such as kekulization and cycle replacements.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations