The Treebank of Vedic Sanskrit (2020.lrec-1)

Copied to clipboard

Challenge: Vedic Sanskrit is a morphologically rich ancient Indian language of central importance for linguistic and historical research.
Approach: They introduce the first treebank of Vedic Sanskrit, a morphologically rich ancient Indian language . they describe how sentences are annotated in the Universal Dependencies scheme and which syntactic constructions required special attention.
Outcome: The proposed treebank reflects the development of metrical and prose texts over a period of 600 years.

Similar Papers

Multi-layer Annotation of the Rigveda (L18-1)

Copied to clipboard

Challenge: Using a multi-level annotation, we present a corpus of the R. GVEDA .
Approach: They propose a multi-level annotation of the R . GVEDA, a Sanskrit text composed in the 2. millenium BCE, and a basic argument identification algorithm to supplement missing verb-argument links.
Outcome: The proposed model replaces verb-argument links by LSTM based model . the proposed model is based on a LS-based model to supplement missing verb-al arguments.
A Universal Dependencies Treebank of Ancient Hebrew (2022.lrec-1)

Copied to clipboard

Challenge: Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew.
Approach: They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax.
Outcome: The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax.
SandhiKosh: A Benchmark Corpus for Evaluating Sanskrit Sandhi Tools (L18-1)

Copied to clipboard

Challenge: Several important texts which are of interest to people all over the world were written in Sanskrit.
Approach: They develop a Sanskrit benchmark to evaluate the completeness and accuracy of tools . they use three most prominent tools to evaluate their completeness .
Outcome: The proposed tools have substantial scope for improvement and are available to researchers worldwide.
A Benchmark and Dataset for Post-OCR text correction in Sanskrit (2022.findings-emnlp)

Copied to clipboard

Challenge: Sanskrit is a classical language with 30 million manuscripts available for digitisation . however, it is considered to be low-resource when it comes to available digital resources.
Approach: They propose to use a post-OCR text correction dataset to correct errors from OCR predictions from 30 different books in the Indian subcontinent.
Outcome: The proposed model outperforms OCR models on graphemic and lexical levels and shows that it is more accurate than previous models.
A New Latin Treebank for Universal Dependencies: Charters between Ancient Latin and Romance Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, Latin features the most data and the most treebanks of all the ancient languages of UD .
Approach: They introduce a Latin treebank that follows the Universal Dependencies (UD) annotation standard . they use a translation of the late Latin Charter Treebank 2 (LLCT2) into the UD style .
Outcome: The proposed treebank is based on the Universal Dependencies (UD) annotation standard.
SanskritShala: A Neural Sanskrit NLP Toolkit with Web-Based Interface for Pedagogical and Annotation Purposes (2023.acl-demo)

Copied to clipboard

Challenge: SanskritShala is a neural-based Sanskrit NLP toolkit that is available as a web-based application .
Approach: They propose a neural Sanskrit NLP toolkit that facilitates linguistic analyses for word segmentation, morphological tagging, dependency parsing, and compound type identification.
Outcome: The proposed toolkit reports state-of-the-art performance on benchmark datasets . it is built with easy-to-use interactive data annotation features .
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources (2026.acl-long)

Copied to clipboard

Challenge: Existing reviews focus on a few high-resource languages or embed Indian languages within broad multilingual settings, limiting coverage of low-resourced and culturally diverse varieties.
Approach: They present a unified survey of Indian NLP resources, covering 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks.
Outcome: The proposed survey covers 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks.
Prague Dependency Treebank - Consolidated 1.0 (2020.lrec-1)

Copied to clipboard

Challenge: Using the standard PDT scheme, the Prague Dependency Treebank-Consolidated 1.0 contains 4 different datasets of Czech, uniformly annotated using the standard scheme.
Approach: They present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0, which contains 4 different datasets of Czech, uniformly annnotated using the standard PDT scheme.
Outcome: The Prague Dependency Treebank-Consolidated 1.0 contains 4 datasets of Czech, uniformly annotated using the standard PDT scheme.
Yorùbá Dependency Treebank (YTB) (2020.lrec-1)

Copied to clipboard

Challenge: Low-resource languages present enormous NLP opportunities as well as varying degrees of difficulties.
Approach: They propose to use the Yoruba Bible treebank to apply a new grammar formalism to the language by examining the use of universal dependency annotations.
Outcome: The treebank of hand-annotated parts of the Yoruba Bible provides an avenue for dependency analysis of the language; the application of a new grammar formalism to the language.
Samayik: A Benchmark and Dataset for English-Sanskrit Translation (2024.lrec-main)

Copied to clipboard

Challenge: Existing Sanskrit corpora focus on poetry and offer limited coverage of contemporary written materials.
Approach: They release a dataset of 53,000 parallel English-Sanskrit sentences . they use spoken content that covers contemporary world affairs and interpretations .
Outcome: a new dataset of 53,000 parallel English-Sanskrit sentences is released . the dataset outperforms existing models trained on older classical-era poetry datasets .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations