MOROCO: The Moldavian and Romanian Dialectal Corpus (P19-1)

Copied to clipboard

Challenge: Using the MOldavian and ROmanian Dialectal COrpus, we perform empirical studies on dialect identification tasks.
Approach: They introduce the MOldavian and ROmanian Dialectal COrpus corpus which contains 33564 samples of text collected from the news domain.
Outcome: The proposed model is based on a shallow and deep approach to discriminate between two different languages.

Similar Papers

SaRoCo: Detecting Satire in a Novel Romanian Corpus of News Articles (2021.acl-short)

Copied to clipboard

Challenge: a corpus for satire detection in Romanian news is based on satirical reporting . the goal is to ridicule public figures, politics or contemporary events .
Approach: They propose a corpus for satire detection in Romanian news . they gather 55,608 public news articles from multiple real and satirical sources .
Outcome: The proposed corpus is one of the largest corpora for satire detection regardless of language . it is the only one for the Romanian language, and the results show that it is low on the machine level compared to human level .
Introducing RONEC - the Romanian Named Entity Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Named Entity Corpus is a free, open-source resource that contains annotated named entities in copy-right free text.
Approach: They present RONEC - the Named Entity Corpus for the Romanian language . it contains over 26000 entities in 5000 annotated sentences belonging to 16 classes .
Outcome: The free, open-source resource contains over 26000 entities in 5000 annotated sentences, belonging to 16 distinct classes.
RoDia: A New Dataset for Romanian Dialect Identification from Speech (2024.findings-naacl)

Copied to clipboard

Challenge: a dataset for Romanian dialect identification from speech is released . the dataset includes speech samples from five distinct regions of Romania .
Approach: They propose a dataset for Romanian dialect identification from speech . they propose competitive models to be used as baselines for future research .
Outcome: The first dataset for Romanian dialect identification from speech is released . the top scoring model achieves 59.83% and 62.08%, respectively .
Dialectal and Low Resource Machine Translation for Aromanian (2025.coling-main)

Copied to clipboard

Challenge: Existing training methods for low-resource languages are focused on English or are massively multilingual, but do not consider the particularities of lowresource language.
Approach: They propose a neural machine translation system that can translate between Romanian, English, and Aromanian.
Outcome: The proposed system can translate between Romanian, English, and Aromanian . BLEU scores range from 17 to 32 depending on direction and genre of text .
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language (2025.findings-emnlp)

Copied to clipboard

Challenge: MoRoVoc is the largest dataset for analyzing the regional variation of spoken Romanian . it has more than 93 hours of audio and 88,192 audio samples .
Approach: They propose a multi-target adversarial training framework that incorporates demographic attributes as adversarials for speech models.
Outcome: The proposed model achieves 78.21% accuracy for variation identification of spoken Romanian using gender as an adversarial target.
Resources in Underrepresented Languages: Building a Representative Romanian Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Currently, the corpus has approximately 5,500,000 tokens originating from written text and 100,000 tokens of spoken language.
Approach: They describe the process of creating a large and representative corpus in Romanian, a relatively under-resourced language with unique typological characteristics.
Outcome: The proposed corpus contains 5,500,000 tokens originating from written text and 100,000 tokens of spoken language.
BioRo: The Biomedical Corpus for the Romanian Language (L18-1)

Copied to clipboard

Challenge: Biomedical text mining uses linguistic resources available in English, but for other languages such as Romanian, the access to language resources is not straight-forward.
Approach: They present a biomedical corpus of the Romanian language, which is a valuable linguistic asset for biomedically text mining.
Outcome: The proposed corpus will be made publicly available to the biomedical text mining community . the corpus is a reference corpus for the Romanian language .
A Multilingual Parallel Corpus for Aromanian (2024.lrec-main)

Copied to clipboard

Challenge: Aromanian is an endangered 1 language that currently lacks corpora and electronic resources that can potentially contribute to the preservation of its cultural heritage.
Approach: They propose to create a corpus of Aromanian and equivalent sentence-aligned translations into Romanian, English, and French using orthographic standards.
Outcome: The authors report that the first high-quality corpus of Aromanian is available in the Balkans and is available for download in Romanian, English, and French.
“Vorbești Românește?” A Recipe to Train Powerful Romanian LLMs with English Instructions (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved almost human-like performance on various tasks.
Approach: They are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate and release open-source LLMs tailored for Romanian.
Outcome: The proposed model trains, evaluates and releases open-source models tailored for Romanian.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations