Papers by Marian Simko

7 papers
Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for genderstereotypical reasoning are limited and often limited to overly specific phenomena.
Approach: They propose to use GEST to measure gender-stereotypical reasoning in language models and machine translation systems.
Outcome: The proposed dataset contains 16 gender stereotypes compatible with the English language and 9 Slavic languages.
SlovakBERT: Slovak Masked Language Model (2022.findings-emnlp)

Copied to clipboard

Challenge: SlovakBERT is a new masked language model that is based on a Web-crawled corpus.
Approach: They introduce a new Slovak-only transformers-based language model called SlovkBERT . they evaluate the model on several NLP tasks and establish a benchmark for Slovakia .
Outcome: The proposed model achieves state-of-the-art on several NLP tasks and achieves best results . the proposed model could be used by other Slovak researchers or NLP practitioners .
Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection (2026.eacl-long)

Copied to clipboard

Challenge: Recent advances in multilingual Large Language Models have enabled powerful capabilities for cross-lingual fact-checking.
Approach: They evaluate six open-source multilingual LLMs across 20 languages using a fully multilingual prompting strategy.
Outcome: The proposed model performs better on high-resource languages than on low-resourced ones.
skLEP: A Slovak General Language Understanding Benchmark (2025.findings-acl)

Copied to clipboard

Challenge: skLEP is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding models.
Approach: They introduce a benchmark specifically designed for evaluating Slovak natural language understanding models.
Outcome: The proposed benchmark covers nine tasks that span token-level, sentence-pair, document-level tasks.
Assessing Web Search Credibility and Response Groundedness in Chat Assistants (2026.eacl-long)

Copied to clipboard

Challenge: Using 100 claims across five misinformation-prone topics, we assess GPT-4o, GPT-5, Perplexity, and Qwen Chat.
Approach: They propose a method for evaluating assistants’ web search behavior focusing on source credibility and the groundedness of responses with respect to cited sources.
Outcome: The proposed method focuses on source credibility and the groundedness of responses with respect to cited sources.
Large Language Models for Multilingual Previously Fact-Checked Claim Detection (2025.findings-emnlp)

Copied to clipboard

Challenge: a new study evaluates large language models for multilingual previously fact-checked claim detection . authors assess seven LLMs across 20 languages in monolingual and cross-lingual settings .
Approach: They evaluate large language models for multilingual previously fact-checked claim detection . they find they perform well for high-resource languages, struggle with low-resourced languages .
Outcome: The proposed model performs well for high-resource languages, but struggle with low-resourced languages.
Soft Language Prompts for Language Transfer (2025.naacl-long)

Copied to clipboard

Challenge: Cross-lingual knowledge transfer, especially between high- and low-resource languages, remains challenging in natural language processing.
Approach: They propose to combine language-specific adapters and soft prompts to enhance cross-lingual transfer by parameter-efficient fine-tuning methods.
Outcome: The proposed methods outperform language adapters and soft prompts in 16 languages and 10 low-resource languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations