Material Philology Meets Digital Onomastic Lexicography: The NordiCon Database of Medieval Nordic Personal Names in Continental Sources (2020.lrec-1)
Copied to clipboard
| Challenge: | Sprkbanken Text is an infrastructure containing modern and historical written data. |
| Approach: | They propose to use a database to contain medieval Nordic personal names attested in Continental sources. |
| Outcome: | The proposed database combines formally interlinked onomastic data with digitized versions of the medieval manuscripts from which the data originate and information on the tokens’ context. |
Similar Papers
The Onomastic Repertoire of the Roman d’Alexandre (ORNARE). Designing an Integrated Digital Onomastic Tool for Medieval French Romance (2024.lrec-main)
Copied to clipboard
| Challenge: | The paper presents the first results of the design and implementation of a new digital tool for romance philology: the Onomastic Repertoire for the medieval French romance (12th-15th centuries). |
| Approach: | The paper presents the design and implementation of a digital romance philology tool . it uses a selection of romances from the corpus of the medieval French Roman d'Alexandre . |
| Outcome: | The proposed system was based on the corpus of the medieval French Roman d'Alexandre . it is the first integrated system for the creation of the Onomastic Repertoire of the romaN d’AlexandRE . |
Dr. Livingstone, I presume? Polishing of foreign character identification in literary texts (2022.naacl-srw)
Copied to clipboard
| Challenge: | Current state-of-the-art models that use neural networks can help with character identification in agglutinative languages. |
| Approach: | They propose to use a search for the shortest version of the name to identify the baseform of the character's lemma to align different appearances of the same character in the narrative. |
| Outcome: | The proposed method is the easiest, best performing and resource-independent method. |
NorNE: Annotating Named Entities for Norwegian (2020.lrec-1)
Copied to clipboard
| Challenge: | Using the annotations of the existing treebank, we have created a dataset for named entity recognition for Norwegian. |
| Approach: | They propose to create a manually annotated corpus of named entities for Norwegian . they propose to add named entity annotations to existing treebank . |
| Outcome: | The proposed dataset extends the annotation of the existing Norwegian Dependency Treebank. |
A Diachronic Treebank of Russian Spanning More Than a Thousand Years (2020.lrec-1)
Copied to clipboard
| Challenge: | TOROT is a treebank that spans from the earliest Old Church Slavonic to modern Russian texts. |
| Approach: | They describe a new version of the Troms Old Russian and Old Church Slavonic Treebank . it adds a modern subcorpus to the existing treebank of contemporary standard Russian . they describe the conversion of SynTagRus into a treebank covering every attested stage of Russian and OCS . |
| Outcome: | The TOROT 20200116 treebank covers all attested stages of Russian and OCS . it includes a modern subcorpus that was created by a conversion of the SynTagRus treebank . |
Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0 (2020.lrec-1)
Copied to clipboard
| Challenge: | Diachronic lexical information is increasingly used in historical linguistics and in NLP . etymological resources need to be fine-grained, large-coverage and accurate . |
| Approach: | They propose guidelines to generate etymological lexical resources for each step of the life-cycle of an ethymology . they introduce EtymDB 2.0, an 'etiological database' generated from the Wiktionary . |
| Outcome: | The proposed resources are generated for each step of the life-cycle of an etymological lexicon: creation, update, evaluation, dissemination, and exploitation. |
CLAUSE-ATLAS: A Corpus of Narrative Information to Scale up Computational Literary Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | XIX and XX century English novels annotated automatically contain 41,715 labeled clauses . a new approach to analyze novels based on clauses captures structural patterns within books, as well as qualitative differences between them. |
| Approach: | They propose to use a corpus of XIX and XX century English novels annotated automatically to study stories as sequences of eventive, subjective and contextual information. |
| Outcome: | The proposed method captures structural patterns within books, as well as qualitative differences between them. |
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)
Copied to clipboard
| Challenge: | BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Approach: | They present a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Outcome: | The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages . |
AGILe: The First Lemmatizer for Ancient Greek Inscriptions (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing models for ancient Greek inscriptions are not performant on epigraphic data due to language differences . a lemmatizer for ancient inscription data can enable meaningful generalizations, we show . |
| Approach: | They propose to train an automatic lemmatizer for ancient Greek inscriptions with 80% accuracy . they also show that existing models are not performant on epigraphic data . |
| Outcome: | The proposed model achieves above 80% accuracy on epigraphic data, and makes it available to the community. |
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)
Copied to clipboard
| Challenge: | a new method is proposed to acquire typological evidence from "gold" treebanks for different languages. |
| Approach: | They propose a method for acquiring typological evidence from "gold" treebanks for different languages. |
| Outcome: | The proposed method can shed light on key issues of the linguistic typological literature. |
Named Entity Recognition in Estonian 19th Century Parish Court Records (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 19th century Parish Court records annotated for named entities (NE) in Estonian is a valuable resource for historians, linguists and the public at large. |
| Approach: | They propose to annotate a corpus of Estonian Parish Court records annotated for named entities (NE) and report on named entity recognition experiments using this corpus. |
| Outcome: | The proposed model achieves microaverage F1 score of 93.6, comparable to state-of-the-art NER performance on the contemporary Estonian. |