Papers by Besim Kabashi
A Corpus of German Reddit Exchanges (GeRedE) (2020.lrec-1)
Copied to clipboard
| Challenge: | Reddit is a popular online platform combining social news aggregation, discussion and microblogging. |
| Approach: | They propose a method to filter out German data and further pre-processing steps to find out what is linguistically peculiar in the German data. |
| Outcome: | The proposed method filters out German data and includes metadata and annotation layers. |
Modelling Frequency, Attestation, and Corpus-Based Information with OntoLex-FrAC (2022.coling-1)
Copied to clipboard
| Challenge: | OntoLex-Lemon has become a de facto standard for lexical resources in the web of data. |
| Approach: | This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information. |
| Outcome: | This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information (OntoLx-FrAC) it is intended to complement OntoLemon with the vocabulary to represent major types of information found in or automatically derived from corpora, for applications in both language technology and the language sciences. |
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
| Approach: | They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics. |
| Outcome: | The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
On Modelling Corpus Citations in Computational Lexical Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | TEI and OntoLex deal with corpus citations in lexicons. |
| Approach: | They argue that TEI and OntoLex can be used to model corpus citations in lexicons . they also argue that they should be combined to achieve a more accurate encoding . |
| Outcome: | The proposed approach favours a combination of TEI and OntoLex . the proposed approach is based on a model of an example entry from a legacy dictionary . |
Albanian Part-of-Speech Tagging: Gold Standard and Evaluation (L18-1)
Copied to clipboard
| Challenge: | a corpus of more than 31,000 tokens is used for part-of-speech tagging in Albanian . a large number of multi-word units are difficult to tally, especially when they have articles or particles as their first part. |
| Approach: | They propose a gold standard corpus for Albanian part-of-speech tagging and perform evaluation experiments with different statistical taggers. |
| Outcome: | The proposed corpus can accurately represent the syntagmatic aspects of Albanian . the results show that the standard is accurate on both the full and coarse tagsets . |