Papers by Taja Kuzman

3 papers
The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild (2022.lrec-1)

Copied to clipboard

Challenge: GINCO is a new training dataset for automatic genre identification based on 1,125 crawled Slovenian web documents that consist of 650,000 words.
Approach: They propose to use 1,125 crawled Slovenian web documents to train a new genre classification system based on a GINCO training dataset .
Outcome: The proposed classifiers perform better on the 1,125 crawled Slovenian web documents than the existing models and achieve higher scores on the task.
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation (2024.lrec-main)

Copied to clipboard

Challenge: Using a similar crawling setup, the corpora are comparable across the entire South Slavic language space.
Approach: They propose to collect 13 billion tokens of texts from 26 million documents . they are linguistically annotated with a CLASSLA-Stanza pipeline and enriched with document-level genre information via a Transformer-based multilingual classifier.
Outcome: The corpora are linguistically annotated with the state-of-the-art CLASSLA-Stanza linguistic processing pipeline and enriched with document-level genre information via the Transformer-based multilingual X-GENRE classifier.
Do Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 Languages (2024.lrec-main)

Copied to clipboard

Challenge: Large, curated, web-crawled corpora play a vital role in training language models . however, relatively little attention has been given to the quality of these corporata .
Approach: They compare four of the currently most relevant large, web-crawled corpora across eleven lower-resourced European languages to evaluate their quality.
Outcome: The CC100 corpus achieves the highest scores on the tests in 11 lower-resourced European languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations