| Challenge: | Using the MOldavian and ROmanian Dialectal COrpus, we perform empirical studies on dialect identification tasks. |
| Approach: | They introduce the MOldavian and ROmanian Dialectal COrpus corpus which contains 33564 samples of text collected from the news domain. |
| Outcome: | The proposed model is based on a shallow and deep approach to discriminate between two different languages. |
Similar Papers
SaRoCo: Detecting Satire in a Novel Romanian Corpus of News Articles (2021.acl-short)
Copied to clipboard
| Challenge: | a corpus for satire detection in Romanian news is based on satirical reporting . the goal is to ridicule public figures, politics or contemporary events . |
| Approach: | They propose a corpus for satire detection in Romanian news . they gather 55,608 public news articles from multiple real and satirical sources . |
| Outcome: | The proposed corpus is one of the largest corpora for satire detection regardless of language . it is the only one for the Romanian language, and the results show that it is low on the machine level compared to human level . |
Introducing RONEC - the Romanian Named Entity Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Named Entity Corpus is a free, open-source resource that contains annotated named entities in copy-right free text. |
| Approach: | They present RONEC - the Named Entity Corpus for the Romanian language . it contains over 26000 entities in 5000 annotated sentences belonging to 16 classes . |
| Outcome: | The free, open-source resource contains over 26000 entities in 5000 annotated sentences, belonging to 16 distinct classes. |
RoDia: A New Dataset for Romanian Dialect Identification from Speech (2024.findings-naacl)
Copied to clipboard
| Challenge: | a dataset for Romanian dialect identification from speech is released . the dataset includes speech samples from five distinct regions of Romania . |
| Approach: | They propose a dataset for Romanian dialect identification from speech . they propose competitive models to be used as baselines for future research . |
| Outcome: | The first dataset for Romanian dialect identification from speech is released . the top scoring model achieves 59.83% and 62.08%, respectively . |
Dialectal and Low Resource Machine Translation for Aromanian (2025.coling-main)
Copied to clipboard
| Challenge: | Existing training methods for low-resource languages are focused on English or are massively multilingual, but do not consider the particularities of lowresource language. |
| Approach: | They propose a neural machine translation system that can translate between Romanian, English, and Aromanian. |
| Outcome: | The proposed system can translate between Romanian, English, and Aromanian . BLEU scores range from 17 to 32 depending on direction and genre of text . |
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language (2025.findings-emnlp)
Copied to clipboard
Andrei-Marius Avram, Bănescu Ema-Ioana, Anda-Teodora Robea, Dumitru-Clementin Cercel, Mihaela-Claudia Cercel
| Challenge: | MoRoVoc is the largest dataset for analyzing the regional variation of spoken Romanian . it has more than 93 hours of audio and 88,192 audio samples . |
| Approach: | They propose a multi-target adversarial training framework that incorporates demographic attributes as adversarials for speech models. |
| Outcome: | The proposed model achieves 78.21% accuracy for variation identification of spoken Romanian using gender as an adversarial target. |
Resources in Underrepresented Languages: Building a Representative Romanian Corpus (2020.lrec-1)
Copied to clipboard
Ludmila Midrigan - Ciochina, Victoria Boyd, Lucila Sanchez-Ortega, Diana Malancea_Malac, Doina Midrigan, David P. Corina
| Challenge: | Currently, the corpus has approximately 5,500,000 tokens originating from written text and 100,000 tokens of spoken language. |
| Approach: | They describe the process of creating a large and representative corpus in Romanian, a relatively under-resourced language with unique typological characteristics. |
| Outcome: | The proposed corpus contains 5,500,000 tokens originating from written text and 100,000 tokens of spoken language. |
BioRo: The Biomedical Corpus for the Romanian Language (L18-1)
Copied to clipboard
| Challenge: | Biomedical text mining uses linguistic resources available in English, but for other languages such as Romanian, the access to language resources is not straight-forward. |
| Approach: | They present a biomedical corpus of the Romanian language, which is a valuable linguistic asset for biomedically text mining. |
| Outcome: | The proposed corpus will be made publicly available to the biomedical text mining community . the corpus is a reference corpus for the Romanian language . |
A Multilingual Parallel Corpus for Aromanian (2024.lrec-main)
Copied to clipboard
| Challenge: | Aromanian is an endangered 1 language that currently lacks corpora and electronic resources that can potentially contribute to the preservation of its cultural heritage. |
| Approach: | They propose to create a corpus of Aromanian and equivalent sentence-aligned translations into Romanian, English, and French using orthographic standards. |
| Outcome: | The authors report that the first high-quality corpus of Aromanian is available in the Balkans and is available for download in Romanian, English, and French. |
“Vorbești Românește?” A Recipe to Train Powerful Romanian LLMs with English Instructions (2024.findings-emnlp)
Copied to clipboard
Mihai Masala, Denis Ilie-Ablachim, Alexandru Dima, Dragos Georgian Corlatescu, Miruna-Andreea Zavelca, Ovio Olaru, Simina-Maria Terian, Andrei Terian, Marius Leordeanu, Horia Velicu, Marius Popescu, Mihai Dascalu, Traian Rebedea
| Challenge: | Large Language Models (LLMs) have achieved almost human-like performance on various tasks. |
| Approach: | They are the first to collect and translate a large collection of texts, instructions, and benchmarks and train, evaluate and release open-source LLMs tailored for Romanian. |
| Outcome: | The proposed model trains, evaluates and releases open-source models tailored for Romanian. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |