Papers by Radu Ion
Collection and Annotation of the Romanian Legal Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, the corpus contains more than 140k documents representing the legislative body of Romania. |
| Approach: | They present a Romanian legislative corpus which is a valuable linguistic asset for machine translation systems. |
| Outcome: | The Romanian legislative corpus contains more than 140k documents representing the legislative body of Romania. |
RACAI’s System at PharmaCoNER 2019 (D19-57)
Copied to clipboard
| Challenge: | RACAI researchers develop named entity recognition systems for Romanian language . current system is language-independent and can be improved by using language-dependent resources . |
| Approach: | They propose to train a named entity recognition system for Romanian language . they propose to use a gazetteer-based baseline and a RNN-based NER system . |
| Outcome: | The proposed system is language independent, provided language-dependent resources exist . the proposed system can detect entities with four labels: anatomical parts, disorders, medical procedures and chemical compounds . |
Ensemble Romanian Dependency Parsing with Neural Networks (L18-1)
Copied to clipboard
| Challenge: | SSPR is a Python 3.5 application based on the Microsoft Cognitive Toolkit 2.0 Python API. |
| Approach: | a Python 3.5 application is based on the Microsoft Cognitive Toolkit 2.0 Python API. |
| Outcome: | SSPR outperforms the best individual parser at the CONLL 2017 dependency parsing shared task. |
The MARCELL Legislative Corpus (2020.lrec-1)
Copied to clipboard
Tamás Váradi, Svetla Koeva, Martin Yamalov, Marko Tadić, Bálint Sass, Bartłomiej Nitoń, Maciej Ogrodniczuk, Piotr Pęzik, Verginica Barbu Mititelu, Radu Ion, Elena Irimia, Maria Mitrofan, Vasile Păiș, Dan Tufiș, Radovan Garabík, Simon Krek, Andraz Repar, Matjaž Rihtar, Janez Brank
| Challenge: | MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification. |
| Approach: | They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents . |
| Outcome: | The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents. |