Papers with Bulgarian

19 papers
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)

Copied to clipboard

Challenge: The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains.
Approach: They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains .
Outcome: The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed .
Entity Framing and Role Portrayal in the News (2025.findings-acl)

Copied to clipboard

Challenge: a dataset of news articles containing 22 fine-grained characters is annotated for entity framing and role portrayal . the dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change .
Approach: They propose a multilingual and hierarchical corpus annotated for entity framing and role portrayal in news articles.
Outcome: The proposed dataset includes 1,378 recent news articles in five languages focusing on the Ukraine-Russia War and climate change . the authors report evaluation results on state-of-the-art multilingual transformers and hierarchical zero-shot learning using LLMs at the level of a document, paragraph, and sentence .
Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society (2021.findings-emnlp)

Copied to clipboard

Challenge: a dataset of 16K manually annotated tweets is used to analyze disinformation . the democratic nature of social media has raised questions about the quality and the factuality of the information that is shared on these platforms.
Approach: They use a dataset of manually annotated tweets to analyze COVID-19 disinformation . they show that tweets contain fake cures, rumors, conspiracy theories and xenophobia .
Outcome: The proposed dataset shows that it is useful in monolingual vs. multilingual settings.
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation (2024.lrec-main)

Copied to clipboard

Challenge: Using a similar crawling setup, the corpora are comparable across the entire South Slavic language space.
Approach: They propose to collect 13 billion tokens of texts from 26 million documents . they are linguistically annotated with a CLASSLA-Stanza pipeline and enriched with document-level genre information via a Transformer-based multilingual classifier.
Outcome: The corpora are linguistically annotated with the state-of-the-art CLASSLA-Stanza linguistic processing pipeline and enriched with document-level genre information via the Transformer-based multilingual X-GENRE classifier.
ISO-based Annotated Multilingual Parallel Corpus for Discourse Markers (2022.lrec-1)

Copied to clipboard

Challenge: Discourse markers carry information about the discourse structure and organization, and also signal local dependencies or epistemic stance of speaker.
Approach: They propose an ISO-based annotated multilingual parallel corpus for discourse markers . they propose an annotation scheme for discourse relations with a plug-in to ISO 24617-2 .
Outcome: The proposed language resource is based on an ISO-based annotated multilingual parallel corpus of discourse markers.
A Deep Transfer Learning Method for Cross-Lingual Natural Language Inference (2022.lrec-1)

Copied to clipboard

Challenge: Natural Language Inference (NLI) is a crucial task in AI and natural language processing.
Approach: They propose an effective transfer learning approach for cross-lingual NLI . they perform experiments on English-Hindi language pairs in cross-linguistic setting .
Outcome: The proposed model improves the baseline model by 10% over the state-of-the-art model.
A Parallel WordNet for English, Swedish and Bulgarian (2020.lrec-1)

Copied to clipboard

Challenge: a new WordNet resource for Swedish and Bulgarian is created that is tightly aligned with the Princeton WordNet.
Approach: They propose a WordNet resource for Swedish and Bulgarian that is tightly aligned with Princeton WordNet.
Outcome: The proposed resource is tightly aligned with the Princeton WordNet for Swedish and Bulgarian . the new resource is open-source and in its development used only existing resources.
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)

Copied to clipboard

Challenge: This work presents a corpus manually annotated with named entities for six Slavic languages .
Approach: They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models .
Outcome: The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier.
NLP for preserving Torlak, a vulnerable low-resource Slavic language (2025.coling-main)

Copied to clipboard

Challenge: Torlak is an endangered, low-resource Slavic language with a high degree of areal and inter-speaker variation.
Approach: They aim to improve the prediction of morphosyntactic annotations for this low-resource Slavic language using the fine-tuning of large language models.
Outcome: The proposed models improve the prediction of morphosyntactic annotations for Torlak using fine-tuning of large language models.
The MARCELL Legislative Corpus (2020.lrec-1)

Copied to clipboard

Challenge: MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification.
Approach: They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents .
Outcome: The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents.
Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages (2022.coling-1)

Copied to clipboard

Challenge: a detailed empirical case study of out-of-vocabulary words in modern text is presented . unfamiliar words cause trouble for machine processing or comprehension of text, authors say .
Approach: They propose a detailed empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian . they apply a multi-faceted distributional analysis of the underlying word-formation processes to characterize the residual vocabulary .
Outcome: The proposed method can be used to aid in compositional translation, parsing, language modeling, and other NLP tasks.
bgGLUE: A Bulgarian General Language Understanding Evaluation Benchmark (2023.acl-long)

Copied to clipboard

Challenge: bgGLUE is a benchmark for evaluating language models on natural language understanding (NLU) tasks in Bulgarian.
Approach: They propose to use a benchmark to evaluate language models on NLU tasks in Bulgarian.
Outcome: The proposed model performs well on sequence labeling tasks, but there is room for improvement for tasks that require more complex reasoning.
Evaluating Word Expansion for Multilingual Sentiment Analysis of Parliamentary Speech (2024.lrec-main)

Copied to clipboard

Challenge: Recent efforts to create and format data sets of parliamentary speech material have facilitated cross-lingual comparisons and highlighted the need for methods that are computationally efficient and language-agnostic.
Approach: They propose a word expansion method for sentiment lexicon generation that leverages word embeddings and vector similarity to expand synonym seed lists with domain-specific terms from the speech corpora.
Outcome: The proposed method is compared with other multilingual lexica and is highly sensitive to processing and scoring techniques.
Modeling the Impact of Syntactic Distance and Surprisal on Cross-Slavic Text Comprehension (2022.lrec-1)

Copied to clipboard

Challenge: Using symmetric measures of insertion, deletion and movement of syntactic units, we investigate phonetic and orthographic asymmetries between selected languages.
Approach: They focus on the syntactic variation and measure syntaktic distances between nine Slavic languages using symmetric measures of insertion, deletion and movement of syntak units in parallel sentences of the fable “The North Wind and the Sun”.
Outcome: The proposed measures are validated on spoken and written cloze tests for Slavic native speakers to determine whether variations in syntax lead to slower or impeded intercomprehension of Slav texts.
Mitigating Catastrophic Forgetting in Language Transfer via Model Merging (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models have shown remarkable capabilities, particularly in English, but for less prevalent languages, performance can be significantly lower, making additional adaptation paramount.
Approach: They propose a new adaptation method based on iteratively merging multiple models fine-tuned on a subset of available training data that reduces forgetting while maintaining learning on the target domain.
Outcome: The proposed method outperforms LLAMA-3-8B-based models in German and German while maintaining learning on the target domain.
NarratEX Dataset: Explaining the Dominant Narratives in News Texts (2025.findings-emnlp)

Copied to clipboard

Challenge: a dataset is created to explain the choice of the dominant narrative in a news article . the dataset is intended to address discourse polarization and propaganda detection .
Approach: They propose a dataset for explaining the choice of the dominant narrative in a news article . the dataset is annotated manually with a dominant narrative and sub-narrative labels .
Outcome: The proposed dataset is designed to explain the choice of the dominant narrative in a news article.
SM-FEEL-BG - the First Bulgarian Datasets and Classifiers for Detecting Feelings, Emotions, and Sentiments of Bulgarian Social Media Text (2024.lrec-main)

Copied to clipboard

Challenge: SM-FEEL-BG is the first Bulgarian-language package for emotion detection and sentiment analysis.
Approach: They introduce SM-FEEL-BG, a Bulgarian-language package that contains 6 datasets with Social Media (SM) texts with emotion, feeling, and sentiment labels and 4 classifiers trained on them.
Outcome: The proposed package is the first to be released in Bulgarian and is available for free.
PolyNarrative: A Multilingual, Multilabel, Multi-domain Dataset for Narrative Extraction from News Articles (2025.acl-long)

Copied to clipboard

Challenge: a new dataset of news articles annotated for narratives provides a framework for narrative detection . recurring narratives can propagate with very high velocity across audiences, languages and countries .
Approach: They propose a multilingual dataset annotated for narratives using two-level taxonomies . they define narrative as a recurring, repetitive, overt or implicit claim that promotes a specific interpretation or viewpoint on an ongoing topic .
Outcome: The proposed dataset will foster research in narrative detection and enable new research directions . the authors identify multiple narratives in the same article, and the results are published online .
Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks suffer from semantic drift and context loss, which can lead to misleading performance metrics.
Approach: They propose a fully automated framework to enable translation of large language models . they propose to use universal self-improvement and multi-round ranking methods to improve translation quality .
Outcome: The proposed framework surpasses existing benchmarks in eight languages and improves translation quality across multilingual domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations