Papers by Oliver Hellwig
Multi-layer Annotation of the Rigveda (L18-1)
Copied to clipboard
| Challenge: | Using a multi-level annotation, we present a corpus of the R. GVEDA . |
| Approach: | They propose a multi-level annotation of the R . GVEDA, a Sanskrit text composed in the 2. millenium BCE, and a basic argument identification algorithm to supplement missing verb-argument links. |
| Outcome: | The proposed model replaces verb-argument links by LSTM based model . the proposed model is based on a LS-based model to supplement missing verb-al arguments. |
Sanskrit Word Segmentation Using Character-level Recurrent and Convolutional Neural Networks (D18-1)
Copied to clipboard
| Challenge: | Using end-to-end neural network models, Sanskrit is tokenized by splitting compounds and resolving phonetic merges. |
| Approach: | They propose end-to-end neural network models that tokenize Sanskrit by jointly splitting compounds and resolving phonetic merges. |
| Outcome: | The proposed models outperform the state-of-the-art for the task of splitting compounds and resolving phonetic merges. |
AET: Web-based Adjective Exploration Tool for German (L18-1)
Copied to clipboard
| Challenge: | AET enables research on the modificational behavior of German adjectives and adverbs . currently available online corpus query tools for German do not lend themselves specifically to research on adjectives - e.g., syntactic relationships or morphological properties. |
| Approach: | They propose a web-based corpus query tool that can be used to query German corpus . they extracted modifiers and modifiees from a print media corpus and stored them in a database . |
| Outcome: | The proposed tool can be transferred to other languages and modification phenomena. |
The Treebank of Vedic Sanskrit (2020.lrec-1)
Copied to clipboard
| Challenge: | Vedic Sanskrit is a morphologically rich ancient Indian language of central importance for linguistic and historical research. |
| Approach: | They introduce the first treebank of Vedic Sanskrit, a morphologically rich ancient Indian language . they describe how sentences are annotated in the Universal Dependencies scheme and which syntactic constructions required special attention. |
| Outcome: | The proposed treebank reflects the development of metrical and prose texts over a period of 600 years. |
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Morphologically rich languages are notoriously challenging to process for downstream NLP applications. |
| Approach: | They propose a pretrained model for NLP applications involving the morphologically rich language Sanskrit that outperforms previous models by a considerable margin. |
| Outcome: | The proposed model outperforms tokenized models on established Sanskrit word segmentation tasks and matches the current best lexicon-based model. |