Papers by Niko Schenk
The ACoLi CoNLL Libraries: Beyond Tab-Separated Values (L18-1)
Copied to clipboard
| Challenge: | a new set of Java archives facilitates advanced manipulations of corpora annotated in TSV formats. |
| Approach: | They propose to use Java archives to facilitate advanced manipulations of corpora annotated in TSV formats. |
| Outcome: | The proposed libraries support all members of the CoNLL format family. |
Towards a Linked Open Data Edition of Sumerian Corpora (L18-1)
Copied to clipboard
| Challenge: | Linguistic Linked Open Data (LLOD) is a flourishing line of research in the language resource community . existing LLOD standards and vocabularies are not widely used in this community despite its popularity . |
| Approach: | They propose to use Linguistic Linked Open Data to link a Sumerian corpus with lexical resources . they use a linguistically annotated archive to create a corpus of cuneiform texts . |
| Outcome: | The proposed LLOD framework is used in assyriology, with philological resources underrepresented . the proposed framework is based on a linguistically annotated corpus of Sumerian texts . |
Knowing the Author by the Company His Words Keep (L18-1)
Copied to clipboard
| Challenge: | In traditional linguistics, there exists a famous saying that one should know a word by the company it keeps. |
| Approach: | They propose a method which uses word embeddings to identify pairwise relational features in the context of authorship attribution. |
| Outcome: | The proposed method is based on three literary corpora and shows that word similarity is a key feature in the authorship attribution task. |
How Low is Too Low? A Computational Perspective on Extremely Low-Resource Languages (2021.acl-srw)
Copied to clipboard
| Challenge: | Sumerian is one of the world’s oldest written languages attested from at least the beginning of the 3rd millennium BC. |
| Approach: | They propose to use interpretLR to train attention-based deep learning models in a low-resource language, Sumerian cuneiform, which includes part-of-speech tagging, named entity recognition, and machine translation. |
| Outcome: | The proposed pipeline outperforms the large language model RoBERTa for POS Tagging and NER. |
Towards the First Machine Translation System for Sumerian Transliterations (2020.coling-main)
Copied to clipboard
| Challenge: | Sumerian cuneiform script was invented more than 5,000 years ago and is one of the oldest in history. |
| Approach: | They propose to translate Sumerian texts into English automatically using supervised, phrase-based, and transfer learning techniques. |
| Outcome: | The proposed method accelerates the costly and time-consuming manual translation process and helps researchers better explore the relationships between Sumerian and Mesopotamian culture. |