Challenge: Sprkbanken Text is an infrastructure containing modern and historical written data.
Approach: They propose to use a database to contain medieval Nordic personal names attested in Continental sources.
Outcome: The proposed database combines formally interlinked onomastic data with digitized versions of the medieval manuscripts from which the data originate and information on the tokens’ context.

Similar Papers

The Onomastic Repertoire of the Roman d’Alexandre (ORNARE). Designing an Integrated Digital Onomastic Tool for Medieval French Romance (2024.lrec-main)

Copied to clipboard

Challenge: The paper presents the first results of the design and implementation of a new digital tool for romance philology: the Onomastic Repertoire for the medieval French romance (12th-15th centuries).
Approach: The paper presents the design and implementation of a digital romance philology tool . it uses a selection of romances from the corpus of the medieval French Roman d'Alexandre .
Outcome: The proposed system was based on the corpus of the medieval French Roman d'Alexandre . it is the first integrated system for the creation of the Onomastic Repertoire of the romaN d’AlexandRE .
Dr. Livingstone, I presume? Polishing of foreign character identification in literary texts (2022.naacl-srw)

Copied to clipboard

Challenge: Current state-of-the-art models that use neural networks can help with character identification in agglutinative languages.
Approach: They propose to use a search for the shortest version of the name to identify the baseform of the character's lemma to align different appearances of the same character in the narrative.
Outcome: The proposed method is the easiest, best performing and resource-independent method.
NorNE: Annotating Named Entities for Norwegian (2020.lrec-1)

Copied to clipboard

Challenge: Using the annotations of the existing treebank, we have created a dataset for named entity recognition for Norwegian.
Approach: They propose to create a manually annotated corpus of named entities for Norwegian . they propose to add named entity annotations to existing treebank .
Outcome: The proposed dataset extends the annotation of the existing Norwegian Dependency Treebank.
A Diachronic Treebank of Russian Spanning More Than a Thousand Years (2020.lrec-1)

Copied to clipboard

Challenge: TOROT is a treebank that spans from the earliest Old Church Slavonic to modern Russian texts.
Approach: They describe a new version of the Troms Old Russian and Old Church Slavonic Treebank . it adds a modern subcorpus to the existing treebank of contemporary standard Russian . they describe the conversion of SynTagRus into a treebank covering every attested stage of Russian and OCS .
Outcome: The TOROT 20200116 treebank covers all attested stages of Russian and OCS . it includes a modern subcorpus that was created by a conversion of the SynTagRus treebank .
Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0 (2020.lrec-1)

Copied to clipboard

Challenge: Diachronic lexical information is increasingly used in historical linguistics and in NLP . etymological resources need to be fine-grained, large-coverage and accurate .
Approach: They propose guidelines to generate etymological lexical resources for each step of the life-cycle of an ethymology . they introduce EtymDB 2.0, an 'etiological database' generated from the Wiktionary .
Outcome: The proposed resources are generated for each step of the life-cycle of an etymological lexicon: creation, update, evaluation, dissemination, and exploitation.
CLAUSE-ATLAS: A Corpus of Narrative Information to Scale up Computational Literary Analysis (2024.lrec-main)

Copied to clipboard

Challenge: XIX and XX century English novels annotated automatically contain 41,715 labeled clauses . a new approach to analyze novels based on clauses captures structural patterns within books, as well as qualitative differences between them.
Approach: They propose to use a corpus of XIX and XX century English novels annotated automatically to study stories as sequences of eventive, subjective and contextual information.
Outcome: The proposed method captures structural patterns within books, as well as qualitative differences between them.
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)

Copied to clipboard

Challenge: BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages.
Approach: They present a database of phonological inventory data from 137 ancient and reconstructed languages.
Outcome: The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages .
AGILe: The First Lemmatizer for Ancient Greek Inscriptions (2022.lrec-1)

Copied to clipboard

Challenge: Existing models for ancient Greek inscriptions are not performant on epigraphic data due to language differences . a lemmatizer for ancient inscription data can enable meaningful generalizations, we show .
Approach: They propose to train an automatic lemmatizer for ancient Greek inscriptions with 80% accuracy . they also show that existing models are not performant on epigraphic data .
Outcome: The proposed model achieves above 80% accuracy on epigraphic data, and makes it available to the community.
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)

Copied to clipboard

Challenge: a new method is proposed to acquire typological evidence from "gold" treebanks for different languages.
Approach: They propose a method for acquiring typological evidence from "gold" treebanks for different languages.
Outcome: The proposed method can shed light on key issues of the linguistic typological literature.
Named Entity Recognition in Estonian 19th Century Parish Court Records (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of 19th century Parish Court records annotated for named entities (NE) in Estonian is a valuable resource for historians, linguists and the public at large.
Approach: They propose to annotate a corpus of Estonian Parish Court records annotated for named entities (NE) and report on named entity recognition experiments using this corpus.
Outcome: The proposed model achieves microaverage F1 score of 93.6, comparable to state-of-the-art NER performance on the contemporary Estonian.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations