Papers by Agnieszka Karlinska

2 papers
Using Bibliodata LODification to Create Metadata-Enriched Literary Corpora in Line with FAIR Principles (2024.lrec-main)

Copied to clipboard

Challenge: Literature corpus building is relatively nascent, and standardized procedures for curating literary corpora are not yet developed.
Approach: They propose a workflow for the creation and reuse of literary corpora using a metadata-enriched Polish Novel Corpus from the 19th and 20th centuries.
Outcome: The proposed workflow includes a multi-stage metadata enrichment and verification process and efficient data collection and data sharing according to the FAIR principles and 5- and 7-star data standards.
BAN-PL: A Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl Web Service (2024.lrec-main)

Copied to clipboard

Challenge: a new dataset of offensive social media content for the Polish language is presented to address this gap . access to accurate and non-synthetic datasets of social media is limited for low-resource languages .
Approach: They present a new open dataset of offensive social media content for the Polish language . authors propose to make the dataset publicly available to improve access .
Outcome: The proposed dataset includes 691,662 posts and comments from the Polish Reddit . the authors describe the dataset and apply it to real-life content moderation processes .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations