Papers by Anastassia Shaitarova

5 papers
SpiritRAG: A Q&A System for Religion and Spirituality in the United Nations Archive (2025.emnlp-demos)

Copied to clipboard

Challenge: Religion and spirituality (R/S) are complex and domain-dependent concepts that have long confounded researchers and policymakers.
Approach: They propose an interactive question-answering system based on Retrieval-Augmented Generation (RAG) SpiritRAG allows researchers and policymakers to conduct complex, context-sensitive database searches of large datasets .
Outcome: SpiritRAG is an interactive Q&A system based on Retrieval-Augmented Generation (RAG) built using 7,500 UN resolution documents related to religion and spirituality in the domains of health and education.
Negation typology and general representation models for cross-lingual zero-shot negation scope resolution in Russian, French, and Spanish. (2021.naacl-srw)

Copied to clipboard

Challenge: Negation resolution remains an acute and continuously researched question in Natural Language Processing.
Approach: They propose to use multilingual pre-trained general representation models to detect negation scope in languages without annotated data.
Outcome: The proposed model achieves token-level F1 score between English, Spanish, French, and Russian.
ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords (2025.acl-long)

Copied to clipboard

Challenge: Lexical borrowing is a ubiquitous linguistic phenomenon influenced by geopolitical, societal, and technological factors.
Approach: They propose a novel contrastive dataset comprising sentences with and without loanwords across 10 languages to examine how machine translation and language models process loanword .
Outcome: The proposed dataset shows that state-of-the-art models prefer loanwords over native terms and exhibit varying performance across languages.
Subword Evenness (SuE) as a Predictor of Cross-lingual Transfer to Low-resource Languages (2022.emnlp-main)

Copied to clipboard

Challenge: English is the most natural choice for cross-lingual transfer, but it is often not the best choice for low-resource languages.
Approach: They propose to use pre-trained multilingual models to improve performance in low-resource languages via cross-lingual transfer.
Outcome: The results show that languages written in non-Latin and non-alphabetic scripts are the best choices for improving performance on Masked Language Modelling tasks in a diverse set of 30 low-resource languages.
Resolving Legalese: A Multilingual Exploration of Negation Scope Resolution in Legal Documents (2024.lrec-main)

Copied to clipboard

Challenge: Negation scope resolution is a challenging task for NLP because of the complexity of legal texts and lack of annotated in-domain negation corpora.
Approach: They propose to use annotated court decisions to improve negation scope resolution . they release annotations in german, french, and italian to train models without legal data .
Outcome: The proposed models achieve token-level F1-scores of up to 86.7% in zero-shot and multilingual settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations