Challenge: Oracle bone script (OBS) documents are the oldest continuously-used writing system in the world and are important for linguistic and historical research.
Approach: They construct an information system for OBS to symbolize, serialize, and store OBS data at the character-level using efficient databases and retrieval modules.
Outcome: The proposed system symbolizes, serializes, and stores OBS data at the character-level, based on efficient databases and retrieval modules.

Similar Papers

Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for deciphering ancient Chinese Oracle Bone Script (OBS) treat deciphering as a closed-set image recognition problem, which fails to bridge the "interpretation gap" .
Approach: They propose a vision-language model framework that integrates a VLM and an LLM to automate a reasoning chain of component identification and knowledge retrieval.
Outcome: The proposed framework yields more detailed and precise decipherments compared to baseline methods.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.
ACSE: An Ancient Character Semantic-Aware Embedding for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on pre-Qin documents are insufficient to understand ancient characters . ancient characters have a low level of digitization and training corpora are extremely scarce .
Approach: They propose a semantic-aware embedding for ancient Chinese characters that integrates glyphs and lexicality into modern Chinese semantic space.
Outcome: The proposed model integrates glyph and lexicality of ancient characters and maps them to the modern Chinese semantic space.
A Dataset of Mycenaean Linear B Sequences (2020.lrec-1)

Copied to clipboard

Challenge: a dataset of Mycenaean Linear B sequences is presented . the dataset contains sequences of Mycean words and ideograms according to the rules of the Mycensean Greek language in the Late Bronze Age.
Approach: They propose to collect Mycenaean Linear B sequences from the Mycensean inscriptions . they exploit the structure of the entire language, not just the Mycean vocabulary .
Outcome: The proposed dataset exploits the structure of the entire language, not just the Mycenaean vocabulary, to analyse sequential patterns.
Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slips (2025.coling-main)

Copied to clipboard

Challenge: Using a multi-modal multi-granularity tokenizer, we analyze ancient Chinese scripts . a large proportion of the characters in ancient Chinese are rare or undeciphered .
Approach: They propose a multi-modal multi-granularity tokenizer specifically designed for ancient Chinese scripts.
Outcome: The proposed tokenizer improves on the part-of-speech tagging task on the Chu bamboo slip script.
Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilities (2022.lrec-1)

Copied to clipboard

Challenge: a brahmic script is used to record endangered languages such as Dogri and Bengali for low-resource languages such that do not require visual normalization.
Approach: They propose to extend Brahmic script functionality within the Nisaba library of finite-state script normalization and processing utilities.
Outcome: The proposed extensions extend coverage from the original ten scripts to an additional ten of South Asia and beyond, including some used to record endangered languages such as Dogri.
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)

Copied to clipboard

Challenge: BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages.
Approach: They present a database of phonological inventory data from 137 ancient and reconstructed languages.
Outcome: The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages .
Finite-state script normalization and processing utilities: The Nisaba Brahmic library (2021.eacl-demos)

Copied to clipboard

Challenge: a library for low-level processing of brahmic scripts is available for free.
Approach: They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts.
Outcome: The proposed library supports low-level processing of ten major south Asian Brahmic scripts.
Sequence Models for Document Structure Identification in an Undeciphered Script (2022.emnlp-main)

Copied to clipboard

Challenge: a systematic analysis of “header” signs in proto-Elamite provides new evidence for their existence . experts have hypothesized that headers are a sign which qualifies transactions .
Approach: They provide unsupervised neural and statistical sequence modeling techniques to identify “header” signs in proto-Elamite, an undeciphered script from 3100-2900 BCE.
Outcome: The authors provide new evidence for the existence of "header" signs in proto-Elamite . they examine which features predict their presence and identify correlations between features and other document properties .
AGILe: The First Lemmatizer for Ancient Greek Inscriptions (2022.lrec-1)

Copied to clipboard

Challenge: Existing models for ancient Greek inscriptions are not performant on epigraphic data due to language differences . a lemmatizer for ancient inscription data can enable meaningful generalizations, we show .
Approach: They propose to train an automatic lemmatizer for ancient Greek inscriptions with 80% accuracy . they also show that existing models are not performant on epigraphic data .
Outcome: The proposed model achieves above 80% accuracy on epigraphic data, and makes it available to the community.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations