John Pavlopoulos, Maria Konstantinidou, Isabelle Marthot-Santaniello, Holger Essler, Asimina Paparigopoulou
| Challenge: | a large number of Greek papyri documents can only be dated tentatively or in approximation due to the lack of decisive evidence. |
| Approach: | a new study trains regression models to estimate Greek papyri's date using a dataset of 389 transcriptions . the authors propose a method to estimate the date of 159 Greek pamphlets, which are only the upper limit known . |
| Outcome: | a new study predicts a date for Greek papyri with an average MAE of 54 years and an MSE of 1.17 . the model outperforms image classifiers and other baselines for 159 manuscripts, with only the upper limit known . |
Similar Papers
Handwritten Paleographic Greek Text Recognition: A Century-Based Approach (2022.lrec-1)
Copied to clipboard
| Challenge: | achieving high accuracy HTR results for Greek manuscripts is still a major challenge . Optical character recognition software is notoriously difficult to use for handwritten text . |
| Approach: | They propose to use Greek manuscripts as a source for a new model to assess HTR accuracy. |
| Outcome: | The proposed model can be used to improve the recognition rate of Greek manuscripts. |
Simple Neologism Based Domain Independent Models to Predict Year of Authorship (C18-1)
Copied to clipboard
| Challenge: | Using domain independent models, we date documents based only on neologism usage patterns . nasa models use only 200 input features, compared to state of the art models using 200K features. |
| Approach: | They propose domain independent models to date documents based only on neologism usage patterns. |
| Outcome: | The proposed models can generalize to various domains like News, Fiction, and Non-Fiction with competitive performance. |
Dating Documents using Graph Convolution Networks (P18-1)
Copied to clipboard
| Challenge: | Existing approaches for document dating assume accurate knowledge of document date, but this is not always available for arbitrary documents from the Web. |
| Approach: | They propose a Graph Convolutional Network (GCN) based document dating approach which exploits syntactic and temporal graph structures of document in a principled way. |
| Outcome: | The proposed approach outperforms state-of-the-art models on real-world datasets by 19% absolute accuracy points. |
AGILe: The First Lemmatizer for Ancient Greek Inscriptions (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing models for ancient Greek inscriptions are not performant on epigraphic data due to language differences . a lemmatizer for ancient inscription data can enable meaningful generalizations, we show . |
| Approach: | They propose to train an automatic lemmatizer for ancient Greek inscriptions with 80% accuracy . they also show that existing models are not performant on epigraphic data . |
| Outcome: | The proposed model achieves above 80% accuracy on epigraphic data, and makes it available to the community. |
Restoring ancient text using deep learning: a case study on Greek epigraphy (D19-1)
Copied to clipboard
| Challenge: | illegible parts of ancient texts must be restored by specialists, known as epigraphists, using deep neural networks to recover missing characters from text input. |
| Approach: | They propose a model that recovers missing characters from a damaged text input using deep neural networks. |
| Outcome: | The proposed model achieves a 30.1% character error rate, compared to the 57.3% of human epigraphists. |
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)
Copied to clipboard
| Challenge: | BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Approach: | They present a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Outcome: | The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages . |
An automatic model and Gold Standard for translation alignment of Ancient Greek (2022.lrec-1)
Copied to clipboard
Tariq Yousef, Chiara Palladino, Farnoosh Shamsian, Anise d’Orange Ferreira, Michel Ferreira dos Reis
| Challenge: | Using a manual annotation tool, we evaluated the performance of various automatic translation alignment models for Ancient Greek. |
| Approach: | They propose a fine-tuning strategy that employs unsupervised training with mono- and bilingual texts and supervised training using manually aligned sentences. |
| Outcome: | The proposed model outperforms the standard on language pairs that were not part of the training data. |
Diachronic word embeddings and semantic shifts: a survey (C18-1)
Copied to clipboard
| Challenge: | Existing methods for tracing time-related semantic shifts with word embedding models lack the cohesion, common terminology and shared practices of more established areas of natural language processing. |
| Approach: | They propose several axes along which these methods can be compared and propose a framework for comparison. |
| Outcome: | The proposed methods are compared with existing methods and outline their main challenges and potential applications. |
A Dataset of Mycenaean Linear B Sequences (2020.lrec-1)
Copied to clipboard
| Challenge: | a dataset of Mycenaean Linear B sequences is presented . the dataset contains sequences of Mycean words and ideograms according to the rules of the Mycensean Greek language in the Late Bronze Age. |
| Approach: | They propose to collect Mycenaean Linear B sequences from the Mycensean inscriptions . they exploit the structure of the entire language, not just the Mycean vocabulary . |
| Outcome: | The proposed dataset exploits the structure of the entire language, not just the Mycenaean vocabulary, to analyse sequential patterns. |
An Annotated Dataset of Errors in Premodern Greek and Baselines for Detecting Them (2025.findings-naacl)
Copied to clipboard
Creston Brooks, Johannes Haubold, Charlie Cowen-Breen, Jay White, Desmond DeVaul, Frederick Riemenschneider, Karthik R Narasimhan, Barbara Graziosi
| Challenge: | a dataset of real errors in premodern Greek is presented to improve error detection methods . scribal errors are more difficult to detect than print or digitization errors. |
| Approach: | They propose to annotate 1,000 words more likely to contain errors and annotated them as errors or not . they propose to evaluate new error detection methods that outperform other methods . |
| Outcome: | The proposed method outperforms all other methods, improving true positive rate by 5%. |