Papers by Nikola Obreshkov
Natural Language Processing Pipeline to Annotate Bulgarian Legislative Documents (2020.lrec-1)
Copied to clipboard
| Challenge: | The Bulgarian MARCELL corpus consists of 25,283 documents, which are classified into eleven types. |
| Approach: | They present the Bulgarian MARCELL corpus, part of a newly developed multilingual corpus representing the national legislation in seven European countries. |
| Outcome: | The proposed corpus represents the national legislation in seven European countries and the NLP pipeline that turns the web crawled data into structured, linguistically annotated dataset. |