Spanish Datasets for Sensitive Entity Detection in the Legal Domain (2022.lrec-1)

Copied to clipboard

Challenge: The de-identification of sensible data is essential for data sharing and reuse, both for research and commercial purposes.
Approach: They propose to use four datasets annotated for named entity detection in Spanish to fine-tune models for the task of named entity-detection.
Outcome: The proposed model is based on four datasets annotated for named entity detection in Spanish with an estimated error rate of 14%.

Similar Papers

Sensitive Data Detection and Classification in Spanish Clinical Text: Experiments with BERT (2020.lrec-1)

Copied to clipboard

Challenge: Massive digital data processing can endanger personal data privacy . anonymisation involves removing or replacing sensitive information from data .
Approach: They propose to use a BERT-based sequence labelling model to conduct an experiment on clinical datasets in Spanish.
Outcome: The proposed model outperforms existing models on clinical datasets in Spanish and shows that it is highly competitive with other models.
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition (2026.findings-eacl)

Copied to clipboard

Challenge: Named entity recognition (NER) is the task of identifying tokens that belong to a predefined set of classes such as "person" or "location"
Approach: They propose a dataset-creation pipeline that scales the teacher-student paradigm to 91 languages and 25 scripts.
Outcome: The proposed model achieves comparable or improved performance in English, Thai, and Swahili despite being trained on 19x less data than strong baselines.
MultiLeg: Dataset for Text Sanitisation in Less-resourced Languages (2024.lrec-main)

Copied to clipboard

Challenge: Text sanitization is the task of detecting and removing personal information from the text.
Approach: They propose a dataset for multilingual named entities that can be used for text sanitization.
Outcome: The proposed dataset is available in 8 languages and contains 3082 parallel text segments for each language.
KIND: an Italian Multi-Domain Dataset for Named Entity Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Named-entity recognition is a task that uses named entities to classify texts . annotated data are time and money consuming, since they need to be created by experts of the domain of the annotation that is going to be done .
Approach: They present an Italian dataset for Named-entity recognition with manual annotations and a semi-automatically annotated part.
Outcome: The proposed dataset covers different styles and language uses, and is the largest in Italy.
A Dataset of German Legal Documents for Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: a dataset developed for Named Entity Recognition in German federal court decisions is available under a CC-BY 4.0 license.
Approach: They describe a dataset developed for Named Entity Recognition in German federal court decisions.
Outcome: The proposed dataset was developed for training an NER service for German legal documents in the EU project Lynx.
MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation) (2022.findings-naacl)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a process of identifying named entities in unstructured texts and classifying them through specific semantic categories.
Approach: They propose a method for automatically producing NER annotations and introduce a manually-annotated test set.
Outcome: The proposed method covers 10 languages, 15 NER categories and 2 textual genres and a manually-annotated test set.
Abusive language in Spanish children and young teenager’s conversations: data preparation and short text classification with contextual word embeddings (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on how to automatically detect abusive short texts are gaining interest in the natural language processing community.
Approach: They propose to use a contextual word embedding model to automatically detect abusive short texts for Spanish language.
Outcome: The proposed model outperforms classical methods in the detection of abusive short texts for the spanish language.
The ApposCorpus: a new multilingual, multi-domain dataset for factual appositive generation (2020.coling-main)

Copied to clipboard

Challenge: appositives are phrases that appear next to a noun phrase and serve an explicative function.
Approach: They propose a more realistic end-to-end definition of appositive generation with a dataset that spans four languages and two entity types.
Outcome: The proposed model is non-trivial and leaves plenty of room for improvement.
Multilingual Entity and Relation Extraction Dataset and Model (2021.eacl-main)

Copied to clipboard

Challenge: HERBERTa is a pipeline for a multilingual task involving two separate BERT models.
Approach: They propose a dataset and a model that combines two independently pretrained BERT models for a multilingual setting to approach the task of Joint Entity and Relation Extraction.
Outcome: The proposed dataset achieves micro F1 81.49 for English on the SMiLER dataset . the proposed pipeline is close to the current SOTA on CoNLL, SpERT .
PPORTAL_ner: An Annotated Corpus of Portuguese Literary Entities (2024.lrec-main)

Copied to clipboard

Challenge: Annotated corpus of 25 literary texts provides a rich set of annotations for Named Entity Recognition models.
Approach: They propose an annotation dataset that simplifies the development of Named Entity Recognition models for Portuguese literary texts.
Outcome: The proposed dataset simplifies the development of Named Entity Recognition models for Portuguese literary works.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations