Papers by Oliver Hellwig

5 papers
Multi-layer Annotation of the Rigveda (L18-1)

Copied to clipboard

Challenge: Using a multi-level annotation, we present a corpus of the R. GVEDA .
Approach: They propose a multi-level annotation of the R . GVEDA, a Sanskrit text composed in the 2. millenium BCE, and a basic argument identification algorithm to supplement missing verb-argument links.
Outcome: The proposed model replaces verb-argument links by LSTM based model . the proposed model is based on a LS-based model to supplement missing verb-al arguments.
Sanskrit Word Segmentation Using Character-level Recurrent and Convolutional Neural Networks (D18-1)

Copied to clipboard

Challenge: Using end-to-end neural network models, Sanskrit is tokenized by splitting compounds and resolving phonetic merges.
Approach: They propose end-to-end neural network models that tokenize Sanskrit by jointly splitting compounds and resolving phonetic merges.
Outcome: The proposed models outperform the state-of-the-art for the task of splitting compounds and resolving phonetic merges.
AET: Web-based Adjective Exploration Tool for German (L18-1)

Copied to clipboard

Challenge: AET enables research on the modificational behavior of German adjectives and adverbs . currently available online corpus query tools for German do not lend themselves specifically to research on adjectives - e.g., syntactic relationships or morphological properties.
Approach: They propose a web-based corpus query tool that can be used to query German corpus . they extracted modifiers and modifiees from a print media corpus and stored them in a database .
Outcome: The proposed tool can be transferred to other languages and modification phenomena.
The Treebank of Vedic Sanskrit (2020.lrec-1)

Copied to clipboard

Challenge: Vedic Sanskrit is a morphologically rich ancient Indian language of central importance for linguistic and historical research.
Approach: They introduce the first treebank of Vedic Sanskrit, a morphologically rich ancient Indian language . they describe how sentences are annotated in the Universal Dependencies scheme and which syntactic constructions required special attention.
Outcome: The proposed treebank reflects the development of metrical and prose texts over a period of 600 years.
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks (2024.findings-emnlp)

Copied to clipboard

Challenge: Morphologically rich languages are notoriously challenging to process for downstream NLP applications.
Approach: They propose a pretrained model for NLP applications involving the morphologically rich language Sanskrit that outperforms previous models by a considerable margin.
Outcome: The proposed model outperforms tokenized models on established Sanskrit word segmentation tasks and matches the current best lexicon-based model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations