Challenge: Currently, there is no way to find 'by-catch', single finds of a different type, in the metadata of excavation reports.
Approach: They propose to train NER classifiers on Dutch excavation reports to help archaeologists find structured information in archaic documents.
Outcome: The proposed dataset contains 31k annotations between six entity types (artefact, time period, place, context, species & material).

Similar Papers

MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation) (2022.findings-naacl)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a process of identifying named entities in unstructured texts and classifying them through specific semantic categories.
Approach: They propose a method for automatically producing NER annotations and introduce a manually-annotated test set.
Outcome: The proposed method covers 10 languages, 15 NER categories and 2 textual genres and a manually-annotated test set.
Towards a Standardized Dataset on Indonesian Named Entity Recognition (2020.aacl-srw)

Copied to clipboard

Challenge: Named entity recognition (NER) tasks in the Indonesian language are still lacking data for the majority of languages, including Indonesian.
Approach: They re-annotated an open dataset with 2,000 sentences and compared the results with a bidirectional long short-term memory and conditional random field approach.
Outcome: The proposed approach improved the prediction score and consistent organization tag for the Indonesian language.
Named Entity Recognition for Entity Linking: What Works and What’s Next (2021.findings-emnlp)

Copied to clipboard

Challenge: Entity Linking (EL) systems have achieved impressive results on standard benchmarks thanks to the contextualized representations provided by recent pretrained language models.
Approach: They propose to exploit Named Entity Recognition (NER) to narrow the gap between EL systems trained on high and low amounts of labeled data.
Outcome: The proposed model can be exploited to narrow the gap between EL systems trained on high and low amounts of labeled data.
Comparing Annotated Datasets for Named Entity Recognition in English Literature (2022.lrec-1)

Copied to clipboard

Challenge: Generally speaking, the majority of NER tools struggle to perform well when the entities in the text contain specific characteristics.
Approach: They analysed two existing annotated datasets and two additional gold standard datasets to evaluate the performance of two NER tools.
Outcome: The results show that the performance of two NER tools varies significantly depending on the gold standard used for the individual evaluations.
Embeddings for Named Entity Recognition in Geoscience Portuguese Literature (2020.lrec-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a task within the field of Natural Language Processing that deals with the identification and categorization of Named entities (NEs) in a given text.
Approach: They propose to use vector and tensor embeddings to train Portuguese Named Entity Recognition (NER) in the Geology domain.
Outcome: The proposed model achieves state-of-the-art for the Portuguese Geology domain with one of its embeddings.
NERetrieve: Dataset for Next Generation Named Entity Recognition and Retrieval (2023.findings-emnlp)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a widely adopted NLP task . authors present three variants of NER task, with dataset to support them .
Approach: They propose three variants of the NER task, together with a dataset to support them . they propose a move towards more fine-grained entities and zero-shot recognition .
Outcome: The proposed model matches or surpasses existing models in NER tasks . the proposed model is based on a large, silver-annotated corpus of 4 million paragraphs .
What Matters for Neural Cross-Lingual Named Entity Recognition: An Empirical Analysis (D19-1)

Copied to clipboard

Challenge: Named entity recognition models are challenging for languages with little training data.
Approach: They propose a simple and efficient neural architecture for cross-lingual named entity recognition models.
Outcome: The proposed model achieves competitive performance with the state-of-the-art on two transferable factors: sequential order and multilingual embedding.
WikiNEuRal: Combined Neural and Knowledge-based Silver Data Creation for Multilingual NER (2021.findings-emnlp)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a key intermediate task in NLP.
Approach: They propose a method which uses knowledge-based approaches and neural models to produce high-quality training corpora for NER.
Outcome: The proposed method improves on standard benchmarks and yields significant improvements up to 6 span-based F1-score points over previous state-of-the-art systems for data creation.
Reconstructing NER Corpora: a Case Study on Bulgarian (2020.lrec-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) and Named Enel Linking (NEL) are two related tasks that are under-resourced for the Slavic languages.
Approach: They propose to use deep learning methods to improve a Named Entity Recognition corpus and to predict and annotate new types in a test corpus.
Outcome: The proposed model improves a type-based Named Entity Recognition (NER) training corpus and predicts and annotates new types in a test corpus.
MasakhaNER: Named Entity Recognition for African Languages (2021.tacl-1)

Copied to clipboard

Challenge: (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results.
Approach: They propose to create a dataset for named entity recognition (NER) in ten African languages.
Outcome: The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations