Papers by Iztok Kosem
Gigafida 2.0: The Reference Corpus of Written Standard Slovene (2020.lrec-1)
Copied to clipboard
Simon Krek, Špela Arhar Holdt, Tomaž Erjavec, Jaka Čibej, Andraz Repar, Polona Gantar, Nikola Ljubešić, Iztok Kosem, Kaja Dobrovoljc
| Challenge: | Gigafida reference corpus of Slovene is updated with new material and tools . focus of upgrade was on transformation from general reference corp to standard reference corp . |
| Approach: | We present a new version of the Gigafida reference corpus of Slovene . the upgrade includes new material and better tools for annotating it . |
| Outcome: | The new version of the Gigafida reference corpus of Slovene is described . the whole Gigido corpus was deduplicated for the first time . |
Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language Learning (2020.lrec-1)
Copied to clipboard
Lionel Nicolas, Verena Lyding, Claudia Borg, Corina Forascu, Karën Fort, Katerina Zdravkova, Iztok Kosem, Jaka Čibej, Špela Arhar Holdt, Alice Millour, Alexander König, Christos Rodosthenous, Federico Sangati, Umair ul Hassan, Anisia Katinskaia, Anabela Barreiro, Lavinia Aparaschivei, Yaakov HaCohen-Kerner
| Challenge: | Lack of wide-coverage and high-quality LRs is a longstanding issue in natural language processing (NLP) however, there are no large initiatives of similar scale for creating new LR or improving existing ones. |
| Approach: | They propose a generic approach to combine implicit crowdsourcing and language learning to mass-produce language resources (LRs) they describe its core paradigm that consists in pairing specific types of LRs with specific exercises . |
| Outcome: | The proposed approach can be used in several learning scenarios to produce a multitude of NLP resources and alleviate the long-standing issue of the lack of LRs. |
SENTA: Sentence Simplification System for Slovene (2024.lrec-main)
Copied to clipboard
| Challenge: | Sentence simplification involves converting complex sentences into more accessible forms while preserving their meaning and context. |
| Approach: | They propose a system for sentence simplification in Slovene that uses a neural classifier to identify sentences that need simplification and a large Slovenen language model to refine sentences into a simpler form. |
| Outcome: | The proposed system achieves an excellent SARI score of 41 for a large Slovene language model based on T5 architecture . it is integrated into a freely accessible, user-friendly user interface, offering a valuable service to less-fluent Slovenen users. |
Towards an Ideal Tool for Learner Error Annotation (2024.lrec-main)
Copied to clipboard
| Challenge: | 'correction annotation' is a technique that has been used for many years to correct errors in learner corpora. |
| Approach: | They propose to use SVALA to annotate and analyse corrections in learner corpora using a parallel aligned approach to visualisation and annotation. |
| Outcome: | The proposed tool supports multiple annotation systems, localisation into other languages, and the development of more complex annotation systems. |