Czech Historical Named Entity Corpus v 1.0 (2020.lrec-1)

Copied to clipboard

Challenge: a lack of annotated historical data for named entity recognition is an obstacle to research in this area.
Approach: They propose to create an annotated corpus for named entity recognition in historical documents . they define domain-specific named entity types and create an annotation manual .
Outcome: The proposed corpus is available for research and is available to download . it is the first annotated historical corpus for named entity recognition (NER)

Similar Papers

People and Places of Historical Europe: Bootstrapping Annotation Pipeline and a New Corpus of Named Entities in Late Medieval Texts (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained named entity recognition models are inaccurate on modern corpora due to differences in language OCR errors.
Approach: They develop a named entity recognition (NER) corpus of 3.6M sentences from medieval charters written mainly in Czech, Latin, and German.
Outcome: The proposed model achieves entity-level Precision of 72.81–93.98% with 58.14–81.77% Recall on a manually-annotated test dataset.
Named Entity Recognition in Estonian 19th Century Parish Court Records (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of 19th century Parish Court records annotated for named entities (NE) in Estonian is a valuable resource for historians, linguists and the public at large.
Approach: They propose to annotate a corpus of Estonian Parish Court records annotated for named entities (NE) and report on named entity recognition experiments using this corpus.
Outcome: The proposed model achieves microaverage F1 score of 93.6, comparable to state-of-the-art NER performance on the contemporary Estonian.
UkraiNER: A New Corpus and Annotation Scheme towards Comprehensive Entity Recognition (2024.lrec-main)

Copied to clipboard

Challenge: Named entity recognition excludes nested, discontinuous, non-named entities in practice . despite attempts to broaden their coverage, the most restrictive variant of NER remains the default .
Approach: They propose a new annotation scheme that offers higher comprehensiveness while preserving simplicity.
Outcome: The proposed scheme offers higher comprehensiveness while preserving simplicity . it also includes an annotation tool to implement the scheme on the corpus UkraiNER .
A Data-driven Approach to Named Entity Recognition for Early Modern French (2022.coling-1)

Copied to clipboard

Challenge: Named entity recognition is an important task in natural language processing.
Approach: They propose to use a data-driven approach to identify historical French with fine-grained annotations instead of a specialised architecture to tackle particularities.
Outcome: The proposed corpus is larger than the most popular NER evaluation corpora for both Contemporary English and French.
A Broad-coverage Corpus for Finnish Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a fundamental task in natural language processing (NLP).
Approach: They propose to annotate Finnish named entity names using a new corpus built on the Universal Dependencies corpus.
Outcome: The new annotation identifies over 10,000 mentions and maintains compatibility with a previously released single-domain corpus for Finnish NER.
A Dataset for Named Entity Recognition and Entity Linking in Chinese Historical Newspapers (2024.lrec-main)

Copied to clipboard

Challenge: a novel historical Chinese dataset is used for named entity recognition, entity linking and entity relations.
Approach: They propose a historical Chinese dataset for named entity recognition, entity linking, coreference and entity relations . they use Chinese newspapers from 1872 to 1949 and multilingual bibliographic resources from the same period .
Outcome: The proposed dataset covers different styles and language uses, and is the largest historical Chinese NER dataset with manual annotations from this transitional period.
Creating a Dataset for Named Entity Recognition in the Archaeology Domain (2020.lrec-1)

Copied to clipboard

Challenge: Currently, there is no way to find 'by-catch', single finds of a different type, in the metadata of excavation reports.
Approach: They propose to train NER classifiers on Dutch excavation reports to help archaeologists find structured information in archaic documents.
Outcome: The proposed dataset contains 31k annotations between six entity types (artefact, time period, place, context, species & material).
NERetrieve: Dataset for Next Generation Named Entity Recognition and Retrieval (2023.findings-emnlp)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a widely adopted NLP task . authors present three variants of NER task, with dataset to support them .
Approach: They propose three variants of the NER task, together with a dataset to support them . they propose a move towards more fine-grained entities and zero-shot recognition .
Outcome: The proposed model matches or surpasses existing models in NER tasks . the proposed model is based on a large, silver-annotated corpus of 4 million paragraphs .
Reconstructing NER Corpora: a Case Study on Bulgarian (2020.lrec-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) and Named Enel Linking (NEL) are two related tasks that are under-resourced for the Slavic languages.
Approach: They propose to use deep learning methods to improve a Named Entity Recognition corpus and to predict and annotate new types in a test corpus.
Outcome: The proposed model improves a type-based Named Entity Recognition (NER) training corpus and predicts and annotates new types in a test corpus.
M-CNER: A Corpus for Chinese Named Entity Recognition in Multi-Domains (L18-1)

Copied to clipboard

Challenge: NER is one of the most important natural language processing tasks.
Approach: They propose to annotate sentences from human-computer interaction, social media, and e-commerce using two rounds of annotation.
Outcome: The proposed system performs the best on all the data sets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations