Dating Greek Papyri with Text Regression (2023.acl-long)

Copied to clipboard

Challenge: a large number of Greek papyri documents can only be dated tentatively or in approximation due to the lack of decisive evidence.
Approach: a new study trains regression models to estimate Greek papyri's date using a dataset of 389 transcriptions . the authors propose a method to estimate the date of 159 Greek pamphlets, which are only the upper limit known .
Outcome: a new study predicts a date for Greek papyri with an average MAE of 54 years and an MSE of 1.17 . the model outperforms image classifiers and other baselines for 159 manuscripts, with only the upper limit known .

Similar Papers

Handwritten Paleographic Greek Text Recognition: A Century-Based Approach (2022.lrec-1)

Copied to clipboard

Challenge: achieving high accuracy HTR results for Greek manuscripts is still a major challenge . Optical character recognition software is notoriously difficult to use for handwritten text .
Approach: They propose to use Greek manuscripts as a source for a new model to assess HTR accuracy.
Outcome: The proposed model can be used to improve the recognition rate of Greek manuscripts.
Simple Neologism Based Domain Independent Models to Predict Year of Authorship (C18-1)

Copied to clipboard

Challenge: Using domain independent models, we date documents based only on neologism usage patterns . nasa models use only 200 input features, compared to state of the art models using 200K features.
Approach: They propose domain independent models to date documents based only on neologism usage patterns.
Outcome: The proposed models can generalize to various domains like News, Fiction, and Non-Fiction with competitive performance.
Dating Documents using Graph Convolution Networks (P18-1)

Copied to clipboard

Challenge: Existing approaches for document dating assume accurate knowledge of document date, but this is not always available for arbitrary documents from the Web.
Approach: They propose a Graph Convolutional Network (GCN) based document dating approach which exploits syntactic and temporal graph structures of document in a principled way.
Outcome: The proposed approach outperforms state-of-the-art models on real-world datasets by 19% absolute accuracy points.
AGILe: The First Lemmatizer for Ancient Greek Inscriptions (2022.lrec-1)

Copied to clipboard

Challenge: Existing models for ancient Greek inscriptions are not performant on epigraphic data due to language differences . a lemmatizer for ancient inscription data can enable meaningful generalizations, we show .
Approach: They propose to train an automatic lemmatizer for ancient Greek inscriptions with 80% accuracy . they also show that existing models are not performant on epigraphic data .
Outcome: The proposed model achieves above 80% accuracy on epigraphic data, and makes it available to the community.
Restoring ancient text using deep learning: a case study on Greek epigraphy (D19-1)

Copied to clipboard

Challenge: illegible parts of ancient texts must be restored by specialists, known as epigraphists, using deep neural networks to recover missing characters from text input.
Approach: They propose a model that recovers missing characters from a damaged text input using deep neural networks.
Outcome: The proposed model achieves a 30.1% character error rate, compared to the 57.3% of human epigraphists.
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)

Copied to clipboard

Challenge: BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages.
Approach: They present a database of phonological inventory data from 137 ancient and reconstructed languages.
Outcome: The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages .
An automatic model and Gold Standard for translation alignment of Ancient Greek (2022.lrec-1)

Copied to clipboard

Challenge: Using a manual annotation tool, we evaluated the performance of various automatic translation alignment models for Ancient Greek.
Approach: They propose a fine-tuning strategy that employs unsupervised training with mono- and bilingual texts and supervised training using manually aligned sentences.
Outcome: The proposed model outperforms the standard on language pairs that were not part of the training data.
Diachronic word embeddings and semantic shifts: a survey (C18-1)

Copied to clipboard

Challenge: Existing methods for tracing time-related semantic shifts with word embedding models lack the cohesion, common terminology and shared practices of more established areas of natural language processing.
Approach: They propose several axes along which these methods can be compared and propose a framework for comparison.
Outcome: The proposed methods are compared with existing methods and outline their main challenges and potential applications.
A Dataset of Mycenaean Linear B Sequences (2020.lrec-1)

Copied to clipboard

Challenge: a dataset of Mycenaean Linear B sequences is presented . the dataset contains sequences of Mycean words and ideograms according to the rules of the Mycensean Greek language in the Late Bronze Age.
Approach: They propose to collect Mycenaean Linear B sequences from the Mycensean inscriptions . they exploit the structure of the entire language, not just the Mycean vocabulary .
Outcome: The proposed dataset exploits the structure of the entire language, not just the Mycenaean vocabulary, to analyse sequential patterns.
An Annotated Dataset of Errors in Premodern Greek and Baselines for Detecting Them (2025.findings-naacl)

Copied to clipboard

Challenge: a dataset of real errors in premodern Greek is presented to improve error detection methods . scribal errors are more difficult to detect than print or digitization errors.
Approach: They propose to annotate 1,000 words more likely to contain errors and annotated them as errors or not . they propose to evaluate new error detection methods that outperform other methods .
Outcome: The proposed method outperforms all other methods, improving true positive rate by 5%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations