Papers by Thomas Proisl
SoMeWeTa: A Part-of-Speech Tagger for German Social Media and Web Texts (L18-1)
Copied to clipboard
| Challenge: | Off-the-shelf part-of-speech taggers perform poorly on web and social media data . this is due to the many unconventional spelling variants that occur in web and twitter texts and that result in a high proportion of out-of vocabulary words. |
| Approach: | They propose to use TIGER corpus as a part-of-speech tagger to train a German part- of-speak tagger on the web and social media data of the EmpiriST 2015 shared task. |
| Outcome: | The proposed tagger significantly improves on the state-of-the-art for both the web and social media data. |
A Corpus of German Reddit Exchanges (GeRedE) (2020.lrec-1)
Copied to clipboard
| Challenge: | Reddit is a popular online platform combining social news aggregation, discussion and microblogging. |
| Approach: | They propose a method to filter out German data and further pre-processing steps to find out what is linguistically peculiar in the German data. |
| Outcome: | The proposed method filters out German data and includes metadata and annotation layers. |
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
| Approach: | They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics. |
| Outcome: | The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
Albanian Part-of-Speech Tagging: Gold Standard and Evaluation (L18-1)
Copied to clipboard
| Challenge: | a corpus of more than 31,000 tokens is used for part-of-speech tagging in Albanian . a large number of multi-word units are difficult to tally, especially when they have articles or particles as their first part. |
| Approach: | They propose a gold standard corpus for Albanian part-of-speech tagging and perform evaluation experiments with different statistical taggers. |
| Outcome: | The proposed corpus can accurately represent the syntagmatic aspects of Albanian . the results show that the standard is accurate on both the full and coarse tagsets . |
Delta vs. N-Gram Tracing: Evaluating the Robustness of Authorship Attribution Methods (L18-1)
Copied to clipboard
| Challenge: | a novel authorship attribution method is developed for short texts . delta measures are well-established, but N-gram tracing is not robust enough . |
| Approach: | They propose to use delta measures and N-gram tracing to compare short texts . they find they are highly sensitive to the choice of authors and texts in the corpus . |
| Outcome: | The proposed methods are highly sensitive to the selection of authors and texts in the comparison corpus. |