Papers by Agnieszka Karlinska
Using Bibliodata LODification to Create Metadata-Enriched Literary Corpora in Line with FAIR Principles (2024.lrec-main)
Copied to clipboard
| Challenge: | Literature corpus building is relatively nascent, and standardized procedures for curating literary corpora are not yet developed. |
| Approach: | They propose a workflow for the creation and reuse of literary corpora using a metadata-enriched Polish Novel Corpus from the 19th and 20th centuries. |
| Outcome: | The proposed workflow includes a multi-stage metadata enrichment and verification process and efficient data collection and data sharing according to the FAIR principles and 5- and 7-star data standards. |
BAN-PL: A Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl Web Service (2024.lrec-main)
Copied to clipboard
Anna Kolos, Inez Okulska, Kinga Głąbińska, Agnieszka Karlinska, Emilia Wisnios, Paweł Ellerik, Andrzej Prałat
| Challenge: | a new dataset of offensive social media content for the Polish language is presented to address this gap . access to accurate and non-synthetic datasets of social media is limited for low-resource languages . |
| Approach: | They present a new open dataset of offensive social media content for the Polish language . authors propose to make the dataset publicly available to improve access . |
| Outcome: | The proposed dataset includes 691,662 posts and comments from the Polish Reddit . the authors describe the dataset and apply it to real-life content moderation processes . |