Papers by Jakub Kopál
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for detecting social-media texts are limited to the English language and longer texts are not easily recognisable by humans. |
| Approach: | They propose to use a multilingual and multi-platform dataset to compare machine-generated text detection methods in the social-media domain to compare them to human-written texts. |
| Outcome: | The proposed dataset contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs. |
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation (2025.acl-long)
Copied to clipboard
Aneta Zugecova, Dominik Macko, Ivan Srba, Robert Moro, Jakub Kopál, Katarína Marcinčinová, Matúš Mesarčík
| Challenge: | Recent large language models generate disinformation news articles following predefined narratives . personalization and disinformation abilities of LLMs have not been studied . |
| Approach: | They evaluate the personalization and disinformation abilities of large language models . they find personalization reduces the safety-filter activations, thus effectively functioning as a jailbreak . |
| Outcome: | The proposed model generates disinformation news articles in english with the lowest quality of personalization. |
CEAID: Benchmark of Multilingual Machine-Generated Text Detection Methods for Central European Languages (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for machine-generated text detection are mostly focused on English . existing methods are almost unusable for non-English languages, leaving the transferability towards these languages unexplored. |
| Approach: | They propose to use a train-language combination to compare MGT detection methods . they focus on multi-domain, multi-generator, and multilingual evaluation . |
| Outcome: | The proposed methods are the most performant in the Central European languages and resistant against obfuscation. |