Papers by Niko Schenk

5 papers
The ACoLi CoNLL Libraries: Beyond Tab-Separated Values (L18-1)

Copied to clipboard

Challenge: a new set of Java archives facilitates advanced manipulations of corpora annotated in TSV formats.
Approach: They propose to use Java archives to facilitate advanced manipulations of corpora annotated in TSV formats.
Outcome: The proposed libraries support all members of the CoNLL format family.
Towards a Linked Open Data Edition of Sumerian Corpora (L18-1)

Copied to clipboard

Challenge: Linguistic Linked Open Data (LLOD) is a flourishing line of research in the language resource community . existing LLOD standards and vocabularies are not widely used in this community despite its popularity .
Approach: They propose to use Linguistic Linked Open Data to link a Sumerian corpus with lexical resources . they use a linguistically annotated archive to create a corpus of cuneiform texts .
Outcome: The proposed LLOD framework is used in assyriology, with philological resources underrepresented . the proposed framework is based on a linguistically annotated corpus of Sumerian texts .
Knowing the Author by the Company His Words Keep (L18-1)

Copied to clipboard

Challenge: In traditional linguistics, there exists a famous saying that one should know a word by the company it keeps.
Approach: They propose a method which uses word embeddings to identify pairwise relational features in the context of authorship attribution.
Outcome: The proposed method is based on three literary corpora and shows that word similarity is a key feature in the authorship attribution task.
How Low is Too Low? A Computational Perspective on Extremely Low-Resource Languages (2021.acl-srw)

Copied to clipboard

Challenge: Sumerian is one of the world’s oldest written languages attested from at least the beginning of the 3rd millennium BC.
Approach: They propose to use interpretLR to train attention-based deep learning models in a low-resource language, Sumerian cuneiform, which includes part-of-speech tagging, named entity recognition, and machine translation.
Outcome: The proposed pipeline outperforms the large language model RoBERTa for POS Tagging and NER.
Towards the First Machine Translation System for Sumerian Transliterations (2020.coling-main)

Copied to clipboard

Challenge: Sumerian cuneiform script was invented more than 5,000 years ago and is one of the oldest in history.
Approach: They propose to translate Sumerian texts into English automatically using supervised, phrase-based, and transfer learning techniques.
Outcome: The proposed method accelerates the costly and time-consuming manual translation process and helps researchers better explore the relationships between Sumerian and Mesopotamian culture.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations