Papers by Andraz Repar
Gigafida 2.0: The Reference Corpus of Written Standard Slovene (2020.lrec-1)
Copied to clipboard
Simon Krek, Špela Arhar Holdt, Tomaž Erjavec, Jaka Čibej, Andraz Repar, Polona Gantar, Nikola Ljubešić, Iztok Kosem, Kaja Dobrovoljc
| Challenge: | Gigafida reference corpus of Slovene is updated with new material and tools . focus of upgrade was on transformation from general reference corp to standard reference corp . |
| Approach: | We present a new version of the Gigafida reference corpus of Slovene . the upgrade includes new material and better tools for annotating it . |
| Outcome: | The new version of the Gigafida reference corpus of Slovene is described . the whole Gigido corpus was deduplicated for the first time . |
The MARCELL Legislative Corpus (2020.lrec-1)
Copied to clipboard
Tamás Váradi, Svetla Koeva, Martin Yamalov, Marko Tadić, Bálint Sass, Bartłomiej Nitoń, Maciej Ogrodniczuk, Piotr Pęzik, Verginica Barbu Mititelu, Radu Ion, Elena Irimia, Maria Mitrofan, Vasile Păiș, Dan Tufiș, Radovan Garabík, Simon Krek, Andraz Repar, Matjaž Rihtar, Janez Brank
| Challenge: | MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification. |
| Approach: | They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents . |
| Outcome: | The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents. |