Papers by Taja Kuzman
The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild (2022.lrec-1)
Copied to clipboard
| Challenge: | GINCO is a new training dataset for automatic genre identification based on 1,125 crawled Slovenian web documents that consist of 650,000 words. |
| Approach: | They propose to use 1,125 crawled Slovenian web documents to train a new genre classification system based on a GINCO training dataset . |
| Outcome: | The proposed classifiers perform better on the 1,125 crawled Slovenian web documents than the existing models and achieve higher scores on the task. |
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation (2024.lrec-main)
Copied to clipboard
| Challenge: | Using a similar crawling setup, the corpora are comparable across the entire South Slavic language space. |
| Approach: | They propose to collect 13 billion tokens of texts from 26 million documents . they are linguistically annotated with a CLASSLA-Stanza pipeline and enriched with document-level genre information via a Transformer-based multilingual classifier. |
| Outcome: | The corpora are linguistically annotated with the state-of-the-art CLASSLA-Stanza linguistic processing pipeline and enriched with document-level genre information via the Transformer-based multilingual X-GENRE classifier. |
Do Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 Languages (2024.lrec-main)
Copied to clipboard
Rik van Noord, Taja Kuzman, Peter Rupnik, Nikola Ljubešić, Miquel Esplà-Gomis, Gema Ramírez-Sánchez, Antonio Toral
| Challenge: | Large, curated, web-crawled corpora play a vital role in training language models . however, relatively little attention has been given to the quality of these corporata . |
| Approach: | They compare four of the currently most relevant large, web-crawled corpora across eleven lower-resourced European languages to evaluate their quality. |
| Outcome: | The CC100 corpus achieves the highest scores on the tests in 11 lower-resourced European languages. |