Papers by Besim Kabashi

5 papers
A Corpus of German Reddit Exchanges (GeRedE) (2020.lrec-1)

Copied to clipboard

Challenge: Reddit is a popular online platform combining social news aggregation, discussion and microblogging.
Approach: They propose a method to filter out German data and further pre-processing steps to find out what is linguistically peculiar in the German data.
Outcome: The proposed method filters out German data and includes metadata and annotation layers.
Modelling Frequency, Attestation, and Corpus-Based Information with OntoLex-FrAC (2022.coling-1)

Copied to clipboard

Challenge: OntoLex-Lemon has become a de facto standard for lexical resources in the web of data.
Approach: This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information.
Outcome: This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information (OntoLx-FrAC) it is intended to complement OntoLemon with the vocabulary to represent major types of information found in or automatically derived from corpora, for applications in both language technology and the language sciences.
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)

Copied to clipboard

Challenge: EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data.
Approach: They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics.
Outcome: The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data.
On Modelling Corpus Citations in Computational Lexical Resources (2024.lrec-main)

Copied to clipboard

Challenge: TEI and OntoLex deal with corpus citations in lexicons.
Approach: They argue that TEI and OntoLex can be used to model corpus citations in lexicons . they also argue that they should be combined to achieve a more accurate encoding .
Outcome: The proposed approach favours a combination of TEI and OntoLex . the proposed approach is based on a model of an example entry from a legacy dictionary .
Albanian Part-of-Speech Tagging: Gold Standard and Evaluation (L18-1)

Copied to clipboard

Challenge: a corpus of more than 31,000 tokens is used for part-of-speech tagging in Albanian . a large number of multi-word units are difficult to tally, especially when they have articles or particles as their first part.
Approach: They propose a gold standard corpus for Albanian part-of-speech tagging and perform evaluation experiments with different statistical taggers.
Outcome: The proposed corpus can accurately represent the syntagmatic aspects of Albanian . the results show that the standard is accurate on both the full and coarse tagsets .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations