| Challenge: | a lack of annotated historical data for named entity recognition is an obstacle to research in this area. |
| Approach: | They propose to create an annotated corpus for named entity recognition in historical documents . they define domain-specific named entity types and create an annotation manual . |
| Outcome: | The proposed corpus is available for research and is available to download . it is the first annotated historical corpus for named entity recognition (NER) |
Similar Papers
People and Places of Historical Europe: Bootstrapping Annotation Pipeline and a New Corpus of Named Entities in Late Medieval Texts (2023.findings-acl)
Copied to clipboard
| Challenge: | Pre-trained named entity recognition models are inaccurate on modern corpora due to differences in language OCR errors. |
| Approach: | They develop a named entity recognition (NER) corpus of 3.6M sentences from medieval charters written mainly in Czech, Latin, and German. |
| Outcome: | The proposed model achieves entity-level Precision of 72.81–93.98% with 58.14–81.77% Recall on a manually-annotated test dataset. |
Named Entity Recognition in Estonian 19th Century Parish Court Records (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 19th century Parish Court records annotated for named entities (NE) in Estonian is a valuable resource for historians, linguists and the public at large. |
| Approach: | They propose to annotate a corpus of Estonian Parish Court records annotated for named entities (NE) and report on named entity recognition experiments using this corpus. |
| Outcome: | The proposed model achieves microaverage F1 score of 93.6, comparable to state-of-the-art NER performance on the contemporary Estonian. |
UkraiNER: A New Corpus and Annotation Scheme towards Comprehensive Entity Recognition (2024.lrec-main)
Copied to clipboard
| Challenge: | Named entity recognition excludes nested, discontinuous, non-named entities in practice . despite attempts to broaden their coverage, the most restrictive variant of NER remains the default . |
| Approach: | They propose a new annotation scheme that offers higher comprehensiveness while preserving simplicity. |
| Outcome: | The proposed scheme offers higher comprehensiveness while preserving simplicity . it also includes an annotation tool to implement the scheme on the corpus UkraiNER . |
A Data-driven Approach to Named Entity Recognition for Early Modern French (2022.coling-1)
Copied to clipboard
| Challenge: | Named entity recognition is an important task in natural language processing. |
| Approach: | They propose to use a data-driven approach to identify historical French with fine-grained annotations instead of a specialised architecture to tackle particularities. |
| Outcome: | The proposed corpus is larger than the most popular NER evaluation corpora for both Contemporary English and French. |
A Broad-coverage Corpus for Finnish Named Entity Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a fundamental task in natural language processing (NLP). |
| Approach: | They propose to annotate Finnish named entity names using a new corpus built on the Universal Dependencies corpus. |
| Outcome: | The new annotation identifies over 10,000 mentions and maintains compatibility with a previously released single-domain corpus for Finnish NER. |
A Dataset for Named Entity Recognition and Entity Linking in Chinese Historical Newspapers (2024.lrec-main)
Copied to clipboard
| Challenge: | a novel historical Chinese dataset is used for named entity recognition, entity linking and entity relations. |
| Approach: | They propose a historical Chinese dataset for named entity recognition, entity linking, coreference and entity relations . they use Chinese newspapers from 1872 to 1949 and multilingual bibliographic resources from the same period . |
| Outcome: | The proposed dataset covers different styles and language uses, and is the largest historical Chinese NER dataset with manual annotations from this transitional period. |
Creating a Dataset for Named Entity Recognition in the Archaeology Domain (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, there is no way to find 'by-catch', single finds of a different type, in the metadata of excavation reports. |
| Approach: | They propose to train NER classifiers on Dutch excavation reports to help archaeologists find structured information in archaic documents. |
| Outcome: | The proposed dataset contains 31k annotations between six entity types (artefact, time period, place, context, species & material). |
NERetrieve: Dataset for Next Generation Named Entity Recognition and Retrieval (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a widely adopted NLP task . authors present three variants of NER task, with dataset to support them . |
| Approach: | They propose three variants of the NER task, together with a dataset to support them . they propose a move towards more fine-grained entities and zero-shot recognition . |
| Outcome: | The proposed model matches or surpasses existing models in NER tasks . the proposed model is based on a large, silver-annotated corpus of 4 million paragraphs . |
Reconstructing NER Corpora: a Case Study on Bulgarian (2020.lrec-1)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) and Named Enel Linking (NEL) are two related tasks that are under-resourced for the Slavic languages. |
| Approach: | They propose to use deep learning methods to improve a Named Entity Recognition corpus and to predict and annotate new types in a test corpus. |
| Outcome: | The proposed model improves a type-based Named Entity Recognition (NER) training corpus and predicts and annotates new types in a test corpus. |
M-CNER: A Corpus for Chinese Named Entity Recognition in Multi-Domains (L18-1)
Copied to clipboard
| Challenge: | NER is one of the most important natural language processing tasks. |
| Approach: | They propose to annotate sentences from human-computer interaction, social media, and e-commerce using two rounds of annotation. |
| Outcome: | The proposed system performs the best on all the data sets. |