| Challenge: | Vedic Sanskrit is a morphologically rich ancient Indian language of central importance for linguistic and historical research. |
| Approach: | They introduce the first treebank of Vedic Sanskrit, a morphologically rich ancient Indian language . they describe how sentences are annotated in the Universal Dependencies scheme and which syntactic constructions required special attention. |
| Outcome: | The proposed treebank reflects the development of metrical and prose texts over a period of 600 years. |
Similar Papers
Multi-layer Annotation of the Rigveda (L18-1)
Copied to clipboard
| Challenge: | Using a multi-level annotation, we present a corpus of the R. GVEDA . |
| Approach: | They propose a multi-level annotation of the R . GVEDA, a Sanskrit text composed in the 2. millenium BCE, and a basic argument identification algorithm to supplement missing verb-argument links. |
| Outcome: | The proposed model replaces verb-argument links by LSTM based model . the proposed model is based on a LS-based model to supplement missing verb-al arguments. |
A Universal Dependencies Treebank of Ancient Hebrew (2022.lrec-1)
Copied to clipboard
| Challenge: | Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew. |
| Approach: | They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax. |
| Outcome: | The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax. |
SandhiKosh: A Benchmark Corpus for Evaluating Sanskrit Sandhi Tools (L18-1)
Copied to clipboard
| Challenge: | Several important texts which are of interest to people all over the world were written in Sanskrit. |
| Approach: | They develop a Sanskrit benchmark to evaluate the completeness and accuracy of tools . they use three most prominent tools to evaluate their completeness . |
| Outcome: | The proposed tools have substantial scope for improvement and are available to researchers worldwide. |
A Benchmark and Dataset for Post-OCR text correction in Sanskrit (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Sanskrit is a classical language with 30 million manuscripts available for digitisation . however, it is considered to be low-resource when it comes to available digital resources. |
| Approach: | They propose to use a post-OCR text correction dataset to correct errors from OCR predictions from 30 different books in the Indian subcontinent. |
| Outcome: | The proposed model outperforms OCR models on graphemic and lexical levels and shows that it is more accurate than previous models. |
A New Latin Treebank for Universal Dependencies: Charters between Ancient Latin and Romance Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, Latin features the most data and the most treebanks of all the ancient languages of UD . |
| Approach: | They introduce a Latin treebank that follows the Universal Dependencies (UD) annotation standard . they use a translation of the late Latin Charter Treebank 2 (LLCT2) into the UD style . |
| Outcome: | The proposed treebank is based on the Universal Dependencies (UD) annotation standard. |
SanskritShala: A Neural Sanskrit NLP Toolkit with Web-Based Interface for Pedagogical and Annotation Purposes (2023.acl-demo)
Copied to clipboard
| Challenge: | SanskritShala is a neural-based Sanskrit NLP toolkit that is available as a web-based application . |
| Approach: | They propose a neural Sanskrit NLP toolkit that facilitates linguistic analyses for word segmentation, morphological tagging, dependency parsing, and compound type identification. |
| Outcome: | The proposed toolkit reports state-of-the-art performance on benchmark datasets . it is built with easy-to-use interactive data annotation features . |
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources (2026.acl-long)
Copied to clipboard
| Challenge: | Existing reviews focus on a few high-resource languages or embed Indian languages within broad multilingual settings, limiting coverage of low-resourced and culturally diverse varieties. |
| Approach: | They present a unified survey of Indian NLP resources, covering 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks. |
| Outcome: | The proposed survey covers 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks. |
Prague Dependency Treebank - Consolidated 1.0 (2020.lrec-1)
Copied to clipboard
Jan Hajič, Eduard Bejček, Jaroslava Hlavacova, Marie Mikulová, Milan Straka, Jan Štěpánek, Barbora Štěpánková
| Challenge: | Using the standard PDT scheme, the Prague Dependency Treebank-Consolidated 1.0 contains 4 different datasets of Czech, uniformly annotated using the standard scheme. |
| Approach: | They present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0, which contains 4 different datasets of Czech, uniformly annnotated using the standard PDT scheme. |
| Outcome: | The Prague Dependency Treebank-Consolidated 1.0 contains 4 datasets of Czech, uniformly annotated using the standard PDT scheme. |
Yorùbá Dependency Treebank (YTB) (2020.lrec-1)
Copied to clipboard
| Challenge: | Low-resource languages present enormous NLP opportunities as well as varying degrees of difficulties. |
| Approach: | They propose to use the Yoruba Bible treebank to apply a new grammar formalism to the language by examining the use of universal dependency annotations. |
| Outcome: | The treebank of hand-annotated parts of the Yoruba Bible provides an avenue for dependency analysis of the language; the application of a new grammar formalism to the language. |
Samayik: A Benchmark and Dataset for English-Sanskrit Translation (2024.lrec-main)
Copied to clipboard
Ayush Maheshwari, Ashim Gupta, Amrith Krishna, Atul Kumar Singh, Ganesh Ramakrishnan, Anil Kumar Gourishetty, Jitin Singla
| Challenge: | Existing Sanskrit corpora focus on poetry and offer limited coverage of contemporary written materials. |
| Approach: | They release a dataset of 53,000 parallel English-Sanskrit sentences . they use spoken content that covers contemporary world affairs and interpretations . |
| Outcome: | a new dataset of 53,000 parallel English-Sanskrit sentences is released . the dataset outperforms existing models trained on older classical-era poetry datasets . |