Papers with Bulgarian
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)
Copied to clipboard
Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadić, Vanja Štefanec, Maciej Ogrodniczuk, Bartłomiej Nitoń, Piotr Pęzik, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan, Dan Tufiș, Radovan Garabík, Simon Krek, Andraž Repar
| Challenge: | The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains. |
| Approach: | They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains . |
| Outcome: | The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed . |
Entity Framing and Role Portrayal in the News (2025.findings-acl)
Copied to clipboard
Tarek Mahmoud, Zhuohan Xie, Dimitar Iliyanov Dimitrov, Nikolaos Nikolaidis, Purificação Silvano, Roman Yangarber, Shivam Sharma, Elisa Sartori, Nicolas Stefanovitch, Giovanni Da San Martino, Jakub Piskorski, Preslav Nakov
| Challenge: | a dataset of news articles containing 22 fine-grained characters is annotated for entity framing and role portrayal . the dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change . |
| Approach: | They propose a multilingual and hierarchical corpus annotated for entity framing and role portrayal in news articles. |
| Outcome: | The proposed dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change . the authors report evaluation results on state-of-the-art multilingual transformers and hierarchical zero-shot learning using LLMs at the level of a document, paragraph, and sentence . |
Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society (2021.findings-emnlp)
Copied to clipboard
Firoj Alam, Shaden Shaar, Fahim Dalvi, Hassan Sajjad, Alex Nikolov, Hamdy Mubarak, Giovanni Da San Martino, Ahmed Abdelali, Nadir Durrani, Kareem Darwish, Abdulaziz Al-Homaid, Wajdi Zaghouani, Tommaso Caselli, Gijs Danoe, Friso Stolk, Britt Bruntink, Preslav Nakov
| Challenge: | a dataset of 16K manually annotated tweets is used to analyze disinformation . the democratic nature of social media has raised questions about the quality and the factuality of the information that is shared on these platforms. |
| Approach: | They use a dataset of manually annotated tweets to analyze COVID-19 disinformation . they show that tweets contain fake cures, rumors, conspiracy theories and xenophobia . |
| Outcome: | The proposed dataset shows that it is useful in monolingual vs. multilingual settings. |
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation (2024.lrec-main)
Copied to clipboard
| Challenge: | Using a similar crawling setup, the corpora are comparable across the entire South Slavic language space. |
| Approach: | They propose to collect 13 billion tokens of texts from 26 million documents . they are linguistically annotated with a CLASSLA-Stanza pipeline and enriched with document-level genre information via a Transformer-based multilingual classifier. |
| Outcome: | The corpora are linguistically annotated with the state-of-the-art CLASSLA-Stanza linguistic processing pipeline and enriched with document-level genre information via the Transformer-based multilingual X-GENRE classifier. |
ISO-based Annotated Multilingual Parallel Corpus for Discourse Markers (2022.lrec-1)
Copied to clipboard
Purificação Silvano, Mariana Damova, Giedrė Valūnaitė Oleškevičienė, Chaya Liebeskind, Christian Chiarcos, Dimitar Trajanov, Ciprian-Octavian Truică, Elena-Simona Apostol, Anna Baczkowska
| Challenge: | Discourse markers carry information about the discourse structure and organization, and also signal local dependencies or epistemic stance of speaker. |
| Approach: | They propose an ISO-based annotated multilingual parallel corpus for discourse markers . they propose an annotation scheme for discourse relations with a plug-in to ISO 24617-2 . |
| Outcome: | The proposed language resource is based on an ISO-based annotated multilingual parallel corpus of discourse markers. |
A Deep Transfer Learning Method for Cross-Lingual Natural Language Inference (2022.lrec-1)
Copied to clipboard
| Challenge: | Natural Language Inference (NLI) is a crucial task in AI and natural language processing. |
| Approach: | They propose an effective transfer learning approach for cross-lingual NLI . they perform experiments on English-Hindi language pairs in cross-linguistic setting . |
| Outcome: | The proposed model improves the baseline model by 10% over the state-of-the-art model. |
A Parallel WordNet for English, Swedish and Bulgarian (2020.lrec-1)
Copied to clipboard
| Challenge: | a new WordNet resource for Swedish and Bulgarian is created that is tightly aligned with the Princeton WordNet. |
| Approach: | They propose a WordNet resource for Swedish and Bulgarian that is tightly aligned with Princeton WordNet. |
| Outcome: | The proposed resource is tightly aligned with the Princeton WordNet for Swedish and Bulgarian . the new resource is open-source and in its development used only existing resources. |
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | This work presents a corpus manually annotated with named entities for six Slavic languages . |
| Approach: | They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models . |
| Outcome: | The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier. |
NLP for preserving Torlak, a vulnerable low-resource Slavic language (2025.coling-main)
Copied to clipboard
| Challenge: | Torlak is an endangered, low-resource Slavic language with a high degree of areal and inter-speaker variation. |
| Approach: | They aim to improve the prediction of morphosyntactic annotations for this low-resource Slavic language using the fine-tuning of large language models. |
| Outcome: | The proposed models improve the prediction of morphosyntactic annotations for Torlak using fine-tuning of large language models. |
The MARCELL Legislative Corpus (2020.lrec-1)
Copied to clipboard
Tamás Váradi, Svetla Koeva, Martin Yamalov, Marko Tadić, Bálint Sass, Bartłomiej Nitoń, Maciej Ogrodniczuk, Piotr Pęzik, Verginica Barbu Mititelu, Radu Ion, Elena Irimia, Maria Mitrofan, Vasile Păiș, Dan Tufiș, Radovan Garabík, Simon Krek, Andraz Repar, Matjaž Rihtar, Janez Brank
| Challenge: | MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification. |
| Approach: | They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents . |
| Outcome: | The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents. |
Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages (2022.coling-1)
Copied to clipboard
| Challenge: | a detailed empirical case study of out-of-vocabulary words in modern text is presented . unfamiliar words cause trouble for machine processing or comprehension of text, authors say . |
| Approach: | They propose a detailed empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian . they apply a multi-faceted distributional analysis of the underlying word-formation processes to characterize the residual vocabulary . |
| Outcome: | The proposed method can be used to aid in compositional translation, parsing, language modeling, and other NLP tasks. |
bgGLUE: A Bulgarian General Language Understanding Evaluation Benchmark (2023.acl-long)
Copied to clipboard
Momchil Hardalov, Pepa Atanasova, Todor Mihaylov, Galia Angelova, Kiril Simov, Petya Osenova, Veselin Stoyanov, Ivan Koychev, Preslav Nakov, Dragomir Radev
| Challenge: | bgGLUE is a benchmark for evaluating language models on natural language understanding (NLU) tasks in Bulgarian. |
| Approach: | They propose to use a benchmark to evaluate language models on NLU tasks in Bulgarian. |
| Outcome: | The proposed model performs well on sequence labeling tasks, but there is room for improvement for tasks that require more complex reasoning. |
Evaluating Word Expansion for Multilingual Sentiment Analysis of Parliamentary Speech (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent efforts to create and format data sets of parliamentary speech material have facilitated cross-lingual comparisons and highlighted the need for methods that are computationally efficient and language-agnostic. |
| Approach: | They propose a word expansion method for sentiment lexicon generation that leverages word embeddings and vector similarity to expand synonym seed lists with domain-specific terms from the speech corpora. |
| Outcome: | The proposed method is compared with other multilingual lexica and is highly sensitive to processing and scoring techniques. |
Modeling the Impact of Syntactic Distance and Surprisal on Cross-Slavic Text Comprehension (2022.lrec-1)
Copied to clipboard
| Challenge: | Using symmetric measures of insertion, deletion and movement of syntactic units, we investigate phonetic and orthographic asymmetries between selected languages. |
| Approach: | They focus on the syntactic variation and measure syntaktic distances between nine Slavic languages using symmetric measures of insertion, deletion and movement of syntak units in parallel sentences of the fable “The North Wind and the Sun”. |
| Outcome: | The proposed measures are validated on spoken and written cloze tests for Slavic native speakers to determine whether variations in syntax lead to slower or impeded intercomprehension of Slav texts. |
Mitigating Catastrophic Forgetting in Language Transfer via Model Merging (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models have shown remarkable capabilities, particularly in English, but for less prevalent languages, performance can be significantly lower, making additional adaptation paramount. |
| Approach: | They propose a new adaptation method based on iteratively merging multiple models fine-tuned on a subset of available training data that reduces forgetting while maintaining learning on the target domain. |
| Outcome: | The proposed method outperforms LLAMA-3-8B-based models in German and German while maintaining learning on the target domain. |
NarratEX Dataset: Explaining the Dominant Narratives in News Texts (2025.findings-emnlp)
Copied to clipboard
Nuno Guimarães, Purificação Silvano, Ricardo Campos, Alipio Jorge, Ana Filipa Pacheco, Dimitar Iliyanov Dimitrov, Nikolaos Nikolaidis, Roman Yangarber, Elisa Sartori, Nicolas Stefanovitch, Preslav Nakov, Jakub Piskorski, Giovanni Da San Martino
| Challenge: | a dataset is created to explain the choice of the dominant narrative in a news article . the dataset is intended to address discourse polarization and propaganda detection . |
| Approach: | They propose a dataset for explaining the choice of the dominant narrative in a news article . the dataset is annotated manually with a dominant narrative and sub-narrative labels . |
| Outcome: | The proposed dataset is designed to explain the choice of the dominant narrative in a news article. |
SM-FEEL-BG - the First Bulgarian Datasets and Classifiers for Detecting Feelings, Emotions, and Sentiments of Bulgarian Social Media Text (2024.lrec-main)
Copied to clipboard
Irina Temnikova, Iva Marinova, Silvia Gargova, Ruslana Margova, Alexander Komarov, Tsvetelina Stefanova, Veneta Kireva, Dimana Vyatrova, Nevena Grigorova, Yordan Mandevski, Stefan Minkov
| Challenge: | SM-FEEL-BG is the first Bulgarian-language package for emotion detection and sentiment analysis. |
| Approach: | They introduce SM-FEEL-BG, a Bulgarian-language package that contains 6 datasets with Social Media (SM) texts with emotion, feeling, and sentiment labels and 4 classifiers trained on them. |
| Outcome: | The proposed package is the first to be released in Bulgarian and is available for free. |
PolyNarrative: A Multilingual, Multilabel, Multi-domain Dataset for Narrative Extraction from News Articles (2025.acl-long)
Copied to clipboard
Nikolaos Nikolaidis, Nicolas Stefanovitch, Purificação Silvano, Dimitar Iliyanov Dimitrov, Roman Yangarber, Nuno Guimarães, Elisa Sartori, Ion Androutsopoulos, Preslav Nakov, Giovanni Da San Martino, Jakub Piskorski
| Challenge: | a new dataset of news articles annotated for narratives provides a framework for narrative detection . recurring narratives can propagate with very high velocity across audiences, languages and countries . |
| Approach: | They propose a multilingual dataset annotated for narratives using two-level taxonomies . they define narrative as a recurring, repetitive, overt or implicit claim that promotes a specific interpretation or viewpoint on an ongoing topic . |
| Outcome: | The proposed dataset will foster research in narrative detection and enable new research directions . the authors identify multiple narratives in the same article, and the results are published online . |
Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks suffer from semantic drift and context loss, which can lead to misleading performance metrics. |
| Approach: | They propose a fully automated framework to enable translation of large language models . they propose to use universal self-improvement and multi-round ranking methods to improve translation quality . |
| Outcome: | The proposed framework surpasses existing benchmarks in eight languages and improves translation quality across multilingual domains. |