Papers by Andrej Ridzik
skLEP: A Slovak General Language Understanding Benchmark (2025.findings-acl)
Copied to clipboard
Marek Suppa, Andrej Ridzik, Daniel Hládek, Tomáš Javůrek, Viktória Ondrejová, Kristína Sásiková, Martin Tamajka, Marian Simko
| Challenge: | skLEP is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding models. |
| Approach: | They introduce a benchmark specifically designed for evaluating Slovak natural language understanding models. |
| Outcome: | The proposed benchmark covers nine tasks that span token-level, sentence-pair, document-level tasks. |
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation (2026.acl-long)
Copied to clipboard
| Challenge: | Slovak embeddings are core infrastructure for semantic search, retrieval-augmented generation (RAG), clustering, and classification. |
| Approach: | They propose a MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language . they use 31 datasets across 7 task types to evaluate the performance of the models . |
| Outcome: | The proposed model achieves competitive performance with proprietary APIs while remaining locally deployable for RAG . the model is based on 31 datasets across 7 task types and is 4 the depth of existing benchmark for Slovak . |
o-MEGA: Optimized Methods for Explanation Generation and Analysis (2025.emnlp-demos)
Copied to clipboard
| Challenge: | a growing number of transformer-based language models have created challenges for model transparency and trustworthiness. |
| Approach: | They propose a tool to automatically identify the most effective explainable AI methods . they evaluate o-mega on a post-claim matching pipeline using a curated dataset . |
| Outcome: | The proposed tool shows that the most effective explainable AI methods can be implemented in semantic matching tasks. |