GRIT: A Dataset of Group Reference Recognition in Italian (2024.lrec-main)

Copied to clipboard

Challenge: a task of automatically recognizing group references has not yet gained much attention within NLP.
Approach: They propose a large-scale dataset for automatic group reference recognition in italian . they verify the validity of the task using a fine-tuned BERT model .
Outcome: The proposed dataset proves that it can be applied to political text analysis and social media analysis.

Similar Papers

KIND: an Italian Multi-Domain Dataset for Named Entity Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Named-entity recognition is a task that uses named entities to classify texts . annotated data are time and money consuming, since they need to be created by experts of the domain of the annotation that is going to be done .
Approach: They present an Italian dataset for Named-entity recognition with manual annotations and a semi-automatically annotated part.
Outcome: The proposed dataset covers different styles and language uses, and is the largest in Italy.
Fine-Grained Evaluation for Entity Linking (D19-1)

Copied to clipboard

Challenge: Entity Linking (EL) is an Information Extraction task that identifies entity mentions in a text corpus and associates them with an unambiguous identifier in KBs such as Wikipedia, BabelNet, DBpedia, Wikidata and YAGO.
Approach: They propose a fine-grained categorization of different types of entity mentions and links and propose 'fuzzy recall' metric to address the lack of consensus and compare a selection of online EL systems.
Outcome: The proposed task offers a bridge between unstructured text and structured KBs, where EL has applications for semantic search, document classification, relation extraction, and more.
Fine-grained Named Entity Annotations for German Biographic Interviews (2020.lrec-1)

Copied to clipboard

Challenge: a NER annotation scheme is adapted for a corpus of transcripts of biographic interviews with emigrants to German . a dataset of spoken data and teaser tweets from newspaper sites are used to test the NER inventory.
Approach: They propose a fine-grained NER annotation scheme with 30 labels and apply it to German data.
Outcome: The proposed NER annotations can be applied to spoken data and teaser tweets from newspaper sites and achieve good inter-annotator agreement.
ITALIC: An Italian Culture-Aware Natural Language Benchmark (2025.naacl-long)

Copied to clipboard

Challenge: ITALIC is a large-scale benchmark dataset of 10,000 multiple-choice questions designed to evaluate the natural language understanding of the Italian language and culture.
Approach: They propose to use a large-scale benchmark dataset to evaluate the natural language understanding of the Italian language and culture.
Outcome: The ITALIC dataset spans 12 domains and uses 17 state-of-the-art LLMs to assess the natural language understanding of the italian language and culture.
MultiCoNER v2: a Large Multilingual dataset for Fine-grained and Noisy Named Entity Recognition (2023.findings-emnlp)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a core task in Natural Language Processing.
Approach: They present a dataset for fine-grained Named Entity Recognition covering 33 entity classes across 12 languages in monolingual and multilingual settings.
Outcome: The proposed dataset covers 33 entity classes across 12 languages in monolingual and multilingual settings.
Practical, Efficient, and Customizable Active Learning for Named Entity Recognition in the Digital Humanities (N19-1)

Copied to clipboard

Challenge: Scholars in interdisciplinary fields like the Digital Humanities are increasingly interested in semantic annotation of specialized corpora.
Approach: They propose an active learning solution for named entity recognition that maximizes a custom model’s improvement per additional unit of manual annotation.
Outcome: The proposed model reduces required annotation by 20-60% and outperforms a competitive active learning baseline.
CoNLL#: Fine-grained Error Analysis and a Corrected Test Set for CoNLL-03 English (2024.lrec-main)

Copied to clipboard

Challenge: a glass ceiling for named entity recognition systems has been suggested for 2021 . however, the performance of the most popular NER benchmarks has plateaued since then . we investigate what NER models are still struggling with .
Approach: They perform a fine-grained evaluation of the model outputs by adding document annotations to the CoNLL-03 English dataset to identify lingering errors.
Outcome: The proposed model is able to correct errors and guide future work.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
CleanCoNLL: A Nearly Noise-Free Named Entity Recognition Dataset (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models achieve F1-scores comparable to or exceed noise level in CoNLL-03 . current models have significant annotation errors, incompleteness, and inconsistencies in the data .
Approach: They propose to add a layer of entity linking annotation to the CoNLL-03 corpus to correct 7.0% of all labels.
Outcome: The proposed approach corrects 7.0% of all labels in the English CoNLL-03 dataset.
Towards a Gold Standard Corpus for Variable Detection and Linking in Social Science Publications (L18-1)

Copied to clipboard

Challenge: a new corpus for detecting and linking survey variables is being developed . the corpus is multilingual and includes manually curated word and phrase alignments .
Approach: They propose to create a corpus for the evaluation of detecting and linking survey variables in social science publications.
Outcome: The proposed corpus is the first gold standard for the variable detection and linking task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations