| Challenge: | Biomedical text mining uses linguistic resources available in English, but for other languages such as Romanian, the access to language resources is not straight-forward. |
| Approach: | They present a biomedical corpus of the Romanian language, which is a valuable linguistic asset for biomedically text mining. |
| Outcome: | The proposed corpus will be made publicly available to the biomedical text mining community . the corpus is a reference corpus for the Romanian language . |
Similar Papers
The Reference Corpus of the Contemporary Romanian Language (CoRoLa) (L18-1)
Copied to clipboard
| Challenge: | a four-year project focused on the creation of a big corpus for contemporary Romanian language is underway . the corpus is the largest publicly available corpus of contemporary Romania . |
| Approach: | a four-year project is focusing on the creation of a big corpus for Romanian language . the corpus is the largest publicly available corpus of the language based in the country . authors propose to use the corpora as a tool to query and listen to the results . |
| Outcome: | a four-year project has created the largest publicly available corpus of Romanian language . the corpus is the result of a project focused on the creation of 'corola.racai.ro' the written component contains 1,257,752,812 tokens, distributed in several languages . |
Resources in Underrepresented Languages: Building a Representative Romanian Corpus (2020.lrec-1)
Copied to clipboard
Ludmila Midrigan - Ciochina, Victoria Boyd, Lucila Sanchez-Ortega, Diana Malancea_Malac, Doina Midrigan, David P. Corina
| Challenge: | Currently, the corpus has approximately 5,500,000 tokens originating from written text and 100,000 tokens of spoken language. |
| Approach: | They describe the process of creating a large and representative corpus in Romanian, a relatively under-resourced language with unique typological characteristics. |
| Outcome: | The proposed corpus contains 5,500,000 tokens originating from written text and 100,000 tokens of spoken language. |
Collection and Annotation of the Romanian Legal Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, the corpus contains more than 140k documents representing the legislative body of Romania. |
| Approach: | They present a Romanian legislative corpus which is a valuable linguistic asset for machine translation systems. |
| Outcome: | The Romanian legislative corpus contains more than 140k documents representing the legislative body of Romania. |
Parallel Corpora for the Biomedical Domain (L18-1)
Copied to clipboard
| Challenge: | Existing corpora of parallel corporata are being used in the biomedical domain . MT is known to support readers' access to textual documents in a language other than their native language . |
| Approach: | They propose to leverage parallel corpora to implement cross-lingual information retrieval or machine translation tools. |
| Outcome: | The proposed corpus is being used in the biomedical task at the conference on machine translation (WMT'16 and WMT'17) it can be leveraged to provide access to health information in languages other than English. |
A Bird’s-eye View of Language Processing Projects at the Romanian Academy (L18-1)
Copied to clipboard
| Challenge: | a recent article outlines five projects that address contemporary Romanian language . the authors argue that a constant accumulation of human expertise is needed to develop complex projects. |
| Approach: | a new article gives a general overview of five AI language-related projects at the Romanian Academy . they focus on the creation of a contemporary Romanian language text and speech corpus and language related applications . |
| Outcome: | a new article gives an overview of five AI language-related projects at the Romanian Academy . the projects address contemporary Romanian language, as well as language related applications . |
Introducing RONEC - the Romanian Named Entity Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Named Entity Corpus is a free, open-source resource that contains annotated named entities in copy-right free text. |
| Approach: | They present RONEC - the Named Entity Corpus for the Romanian language . it contains over 26000 entities in 5000 annotated sentences belonging to 16 classes . |
| Outcome: | The free, open-source resource contains over 26000 entities in 5000 annotated sentences, belonging to 16 distinct classes. |
MOROCO: The Moldavian and Romanian Dialectal Corpus (P19-1)
Copied to clipboard
| Challenge: | Using the MOldavian and ROmanian Dialectal COrpus, we perform empirical studies on dialect identification tasks. |
| Approach: | They introduce the MOldavian and ROmanian Dialectal COrpus corpus which contains 33564 samples of text collected from the news domain. |
| Outcome: | The proposed model is based on a shallow and deep approach to discriminate between two different languages. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Named Entities in Medical Case Reports: Corpus and Experiments (2020.lrec-1)
Copied to clipboard
| Challenge: | Only very few annotated corpora in the medical domain exist. |
| Approach: | They propose to annotate medical entities in case reports from PubMed Central's open access library. |
| Outcome: | The proposed corpus is the first of its kind to be made available to the scientific community in English. |
BioMegatron: Larger Biomedical Domain Language Model (2020.emnlp-main)
Copied to clipboard
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, Raghav Mani
| Challenge: | Existing studies on domain language models do not study the factors affecting performance on domain languages. |
| Approach: | They empirically evaluate factors that can affect performance on domain language applications . sub-word vocabulary set, model size, pre-training corpus, and domain transfer are important . |
| Outcome: | The results show language models trained on biomedical text perform better on biomedicine benchmarks than those trained on general domain text corpora. |