The Reference Corpus of the Contemporary Romanian Language (CoRoLa) (L18-1)

Copied to clipboard

Challenge: a four-year project focused on the creation of a big corpus for contemporary Romanian language is underway . the corpus is the largest publicly available corpus of contemporary Romania .
Approach: a four-year project is focusing on the creation of a big corpus for Romanian language . the corpus is the largest publicly available corpus of the language based in the country . authors propose to use the corpora as a tool to query and listen to the results .
Outcome: a four-year project has created the largest publicly available corpus of Romanian language . the corpus is the result of a project focused on the creation of 'corola.racai.ro' the written component contains 1,257,752,812 tokens, distributed in several languages .

Similar Papers

BioRo: The Biomedical Corpus for the Romanian Language (L18-1)

Copied to clipboard

Challenge: Biomedical text mining uses linguistic resources available in English, but for other languages such as Romanian, the access to language resources is not straight-forward.
Approach: They present a biomedical corpus of the Romanian language, which is a valuable linguistic asset for biomedically text mining.
Outcome: The proposed corpus will be made publicly available to the biomedical text mining community . the corpus is a reference corpus for the Romanian language .
Resources in Underrepresented Languages: Building a Representative Romanian Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Currently, the corpus has approximately 5,500,000 tokens originating from written text and 100,000 tokens of spoken language.
Approach: They describe the process of creating a large and representative corpus in Romanian, a relatively under-resourced language with unique typological characteristics.
Outcome: The proposed corpus contains 5,500,000 tokens originating from written text and 100,000 tokens of spoken language.
Collection and Annotation of the Romanian Legal Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Currently, the corpus contains more than 140k documents representing the legislative body of Romania.
Approach: They present a Romanian legislative corpus which is a valuable linguistic asset for machine translation systems.
Outcome: The Romanian legislative corpus contains more than 140k documents representing the legislative body of Romania.
RSC: A Romanian Read Speech Corpus for Automatic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Romanian language is under-resourced due to the lack of acoustic and linguistic resources.
Approach: They propose to use a Romanian speech corpus to train automatic speech recognition algorithms based on the spoken hotword detection mechanism.
Outcome: The read speech corpus is a speech recognition system that can perform automatic speech recognition and speech synthesis using state-of-the-art speech recognition toolkit.
Introducing RONEC - the Romanian Named Entity Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Named Entity Corpus is a free, open-source resource that contains annotated named entities in copy-right free text.
Approach: They present RONEC - the Named Entity Corpus for the Romanian language . it contains over 26000 entities in 5000 annotated sentences belonging to 16 classes .
Outcome: The free, open-source resource contains over 26000 entities in 5000 annotated sentences, belonging to 16 distinct classes.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
The MARCELL Legislative Corpus (2020.lrec-1)

Copied to clipboard

Challenge: MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification.
Approach: They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents .
Outcome: The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents.
A Bird’s-eye View of Language Processing Projects at the Romanian Academy (L18-1)

Copied to clipboard

Challenge: a recent article outlines five projects that address contemporary Romanian language . the authors argue that a constant accumulation of human expertise is needed to develop complex projects.
Approach: a new article gives a general overview of five AI language-related projects at the Romanian Academy . they focus on the creation of a contemporary Romanian language text and speech corpus and language related applications .
Outcome: a new article gives an overview of five AI language-related projects at the Romanian Academy . the projects address contemporary Romanian language, as well as language related applications .
A Multilingual Parallel Corpus for Aromanian (2024.lrec-main)

Copied to clipboard

Challenge: Aromanian is an endangered 1 language that currently lacks corpora and electronic resources that can potentially contribute to the preservation of its cultural heritage.
Approach: They propose to create a corpus of Aromanian and equivalent sentence-aligned translations into Romanian, English, and French using orthographic standards.
Outcome: The authors report that the first high-quality corpus of Aromanian is available in the Balkans and is available for download in Romanian, English, and French.
Manually Annotated Corpus of Polish Texts Published between 1830 and 1918 (L18-1)

Copied to clipboard

Challenge: a paper presents a manually annotated corpus of 625,000 tokens of Polish texts . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Approach: The paper presents a manually annotated large historical corpus of Polish . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Outcome: The corpus provides three layers: transliteration, transcription and morphosyntactic annotation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations