Papers by Thomas Proisl

5 papers
SoMeWeTa: A Part-of-Speech Tagger for German Social Media and Web Texts (L18-1)

Copied to clipboard

Challenge: Off-the-shelf part-of-speech taggers perform poorly on web and social media data . this is due to the many unconventional spelling variants that occur in web and twitter texts and that result in a high proportion of out-of vocabulary words.
Approach: They propose to use TIGER corpus as a part-of-speech tagger to train a German part- of-speak tagger on the web and social media data of the EmpiriST 2015 shared task.
Outcome: The proposed tagger significantly improves on the state-of-the-art for both the web and social media data.
A Corpus of German Reddit Exchanges (GeRedE) (2020.lrec-1)

Copied to clipboard

Challenge: Reddit is a popular online platform combining social news aggregation, discussion and microblogging.
Approach: They propose a method to filter out German data and further pre-processing steps to find out what is linguistically peculiar in the German data.
Outcome: The proposed method filters out German data and includes metadata and annotation layers.
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)

Copied to clipboard

Challenge: EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data.
Approach: They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics.
Outcome: The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data.
Albanian Part-of-Speech Tagging: Gold Standard and Evaluation (L18-1)

Copied to clipboard

Challenge: a corpus of more than 31,000 tokens is used for part-of-speech tagging in Albanian . a large number of multi-word units are difficult to tally, especially when they have articles or particles as their first part.
Approach: They propose a gold standard corpus for Albanian part-of-speech tagging and perform evaluation experiments with different statistical taggers.
Outcome: The proposed corpus can accurately represent the syntagmatic aspects of Albanian . the results show that the standard is accurate on both the full and coarse tagsets .
Delta vs. N-Gram Tracing: Evaluating the Robustness of Authorship Attribution Methods (L18-1)

Copied to clipboard

Challenge: a novel authorship attribution method is developed for short texts . delta measures are well-established, but N-gram tracing is not robust enough .
Approach: They propose to use delta measures and N-gram tracing to compare short texts . they find they are highly sensitive to the choice of authors and texts in the corpus .
Outcome: The proposed methods are highly sensitive to the selection of authors and texts in the comparison corpus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations