DaNE: A Named Entity Resource for Danish (2020.lrec-1)

Copied to clipboard

Challenge: a named entity annotation for the Danish Universal Dependencies treebank is the largest publicly available named entity gold annotation.
Approach: They propose a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme DaNE.
Outcome: The proposed annotations improve Danish named entity recognition over a recent cross-lingual approach and over norwegian training set.

Similar Papers

DaN+: Danish Nested Named Entities and Lexical Normalization (2020.coling-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a task of finding entities in text, such as locations, organizations, and persons.
Approach: They propose a multi-domain corpus and annotation guidelines for Danish nested named entities and lexical normalization to support research on cross-lingual cross-domain learning for a less-resourced language.
Outcome: The proposed model outperforms existing models on the Danish Named Entity Recognition task and shows that it is robust to domain shifts and is highly effective on the least canonical data.
Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark (2024.naacl-long)

Copied to clipboard

Challenge: In named entity recognition, the majority of annotation efforts are centered on English, and cross-lingual transfer performance remains brittle.
Approach: They propose to develop gold-standard named entity recognition benchmarks in many languages using a cross-lingual consistent schema.
Outcome: The proposed benchmarks will be released to the public in 2022 . they will provide baselines on in-language and cross-lingual learning settings.
NorNE: Annotating Named Entities for Norwegian (2020.lrec-1)

Copied to clipboard

Challenge: Using the annotations of the existing treebank, we have created a dataset for named entity recognition for Norwegian.
Approach: They propose to create a manually annotated corpus of named entities for Norwegian . they propose to add named entity annotations to existing treebank .
Outcome: The proposed dataset extends the annotation of the existing Norwegian Dependency Treebank.
A Broad-coverage Corpus for Finnish Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a fundamental task in natural language processing (NLP).
Approach: They propose to annotate Finnish named entity names using a new corpus built on the Universal Dependencies corpus.
Outcome: The new annotation identifies over 10,000 mentions and maintains compatibility with a previously released single-domain corpus for Finnish NER.
NERetrieve: Dataset for Next Generation Named Entity Recognition and Retrieval (2023.findings-emnlp)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a widely adopted NLP task . authors present three variants of NER task, with dataset to support them .
Approach: They propose three variants of the NER task, together with a dataset to support them . they propose a move towards more fine-grained entities and zero-shot recognition .
Outcome: The proposed model matches or surpasses existing models in NER tasks . the proposed model is based on a large, silver-annotated corpus of 4 million paragraphs .
CleanCoNLL: A Nearly Noise-Free Named Entity Recognition Dataset (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models achieve F1-scores comparable to or exceed noise level in CoNLL-03 . current models have significant annotation errors, incompleteness, and inconsistencies in the data .
Approach: They propose to add a layer of entity linking annotation to the CoNLL-03 corpus to correct 7.0% of all labels.
Outcome: The proposed approach corrects 7.0% of all labels in the English CoNLL-03 dataset.
A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers (D19-1)

Copied to clipboard

Challenge: Named entity recognition models rely on large amounts of labeled data, making them challenging to extend to new, lower-resource languages.
Approach: They propose a method for bootstrapping named entity recognition models in under-resourced languages . they use cross-lingual transfer learning and targeted annotation of only uncertain entities .
Outcome: The proposed method achieves competitive accuracy with just one-tenth of training data.
Sources of Transfer in Multilingual Named Entity Recognition (2020.acl-main)

Copied to clipboard

Challenge: naive training of named-entity recognition models using annotated data from multiple languages consistently underperforms monolingual models.
Approach: They propose a polyglot named-entity recognition model where one model is trained using annotated data drawn from multiple languages.
Outcome: The proposed model outperforms models trained on monolingual data despite more training data . the proposed model shares many parameters across languages and fine-tunes them to outperFORM monolingual models.
Cross-Lingual Cross-Domain Nested Named Entity Evaluation on English Web Texts (2021.findings-acl)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a key task in Natural Language Processing, but most existing work on NER ignores the recognition of nested entities.
Approach: They propose to annotate five web domains for nested named entities on top of the English Web Treebank (EWT) . they propose to use the English web treebank to perform cross-domain evaluations.
Outcome: The proposed dataset covers five domains and includes transfer results from German and Danish.
Beyond Boundaries: Learning a Universal Entity Taxonomy across Datasets and Languages for Open Named Entity Recognition (2025.coling-main)

Copied to clipboard

Challenge: Current Large Language Models struggle with complex entity taxonomies in open domains and lack NER capabilities.
Approach: They propose a dataset to guide LLMs' generalization in Open NER under a universal entity taxonomy.
Outcome: The proposed model outperforms GPT-4 in 3 out-of-domain benchmarks across 15 datasets and 6 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations