Spanish Datasets for Sensitive Entity Detection in the Legal Domain (2022.lrec-1)
Copied to clipboard
| Challenge: | The de-identification of sensible data is essential for data sharing and reuse, both for research and commercial purposes. |
| Approach: | They propose to use four datasets annotated for named entity detection in Spanish to fine-tune models for the task of named entity-detection. |
| Outcome: | The proposed model is based on four datasets annotated for named entity detection in Spanish with an estimated error rate of 14%. |
Similar Papers
Sensitive Data Detection and Classification in Spanish Clinical Text: Experiments with BERT (2020.lrec-1)
Copied to clipboard
| Challenge: | Massive digital data processing can endanger personal data privacy . anonymisation involves removing or replacing sensitive information from data . |
| Approach: | They propose to use a BERT-based sequence labelling model to conduct an experiment on clinical datasets in Spanish. |
| Outcome: | The proposed model outperforms existing models on clinical datasets in Spanish and shows that it is highly competitive with other models. |
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition (2026.findings-eacl)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is the task of identifying tokens that belong to a predefined set of classes such as "person" or "location" |
| Approach: | They propose a dataset-creation pipeline that scales the teacher-student paradigm to 91 languages and 25 scripts. |
| Outcome: | The proposed model achieves comparable or improved performance in English, Thai, and Swahili despite being trained on 19x less data than strong baselines. |
MultiLeg: Dataset for Text Sanitisation in Less-resourced Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | Text sanitization is the task of detecting and removing personal information from the text. |
| Approach: | They propose a dataset for multilingual named entities that can be used for text sanitization. |
| Outcome: | The proposed dataset is available in 8 languages and contains 3082 parallel text segments for each language. |
KIND: an Italian Multi-Domain Dataset for Named Entity Recognition (2022.lrec-1)
Copied to clipboard
| Challenge: | Named-entity recognition is a task that uses named entities to classify texts . annotated data are time and money consuming, since they need to be created by experts of the domain of the annotation that is going to be done . |
| Approach: | They present an Italian dataset for Named-entity recognition with manual annotations and a semi-automatically annotated part. |
| Outcome: | The proposed dataset covers different styles and language uses, and is the largest in Italy. |
A Dataset of German Legal Documents for Named Entity Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | a dataset developed for Named Entity Recognition in German federal court decisions is available under a CC-BY 4.0 license. |
| Approach: | They describe a dataset developed for Named Entity Recognition in German federal court decisions. |
| Outcome: | The proposed dataset was developed for training an NER service for German legal documents in the EU project Lynx. |
MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation) (2022.findings-naacl)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a process of identifying named entities in unstructured texts and classifying them through specific semantic categories. |
| Approach: | They propose a method for automatically producing NER annotations and introduce a manually-annotated test set. |
| Outcome: | The proposed method covers 10 languages, 15 NER categories and 2 textual genres and a manually-annotated test set. |
Abusive language in Spanish children and young teenager’s conversations: data preparation and short text classification with contextual word embeddings (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on how to automatically detect abusive short texts are gaining interest in the natural language processing community. |
| Approach: | They propose to use a contextual word embedding model to automatically detect abusive short texts for Spanish language. |
| Outcome: | The proposed model outperforms classical methods in the detection of abusive short texts for the spanish language. |
The ApposCorpus: a new multilingual, multi-domain dataset for factual appositive generation (2020.coling-main)
Copied to clipboard
| Challenge: | appositives are phrases that appear next to a noun phrase and serve an explicative function. |
| Approach: | They propose a more realistic end-to-end definition of appositive generation with a dataset that spans four languages and two entity types. |
| Outcome: | The proposed model is non-trivial and leaves plenty of room for improvement. |
Multilingual Entity and Relation Extraction Dataset and Model (2021.eacl-main)
Copied to clipboard
| Challenge: | HERBERTa is a pipeline for a multilingual task involving two separate BERT models. |
| Approach: | They propose a dataset and a model that combines two independently pretrained BERT models for a multilingual setting to approach the task of Joint Entity and Relation Extraction. |
| Outcome: | The proposed dataset achieves micro F1 81.49 for English on the SMiLER dataset . the proposed pipeline is close to the current SOTA on CoNLL, SpERT . |
PPORTAL_ner: An Annotated Corpus of Portuguese Literary Entities (2024.lrec-main)
Copied to clipboard
| Challenge: | Annotated corpus of 25 literary texts provides a rich set of annotations for Named Entity Recognition models. |
| Approach: | They propose an annotation dataset that simplifies the development of Named Entity Recognition models for Portuguese literary texts. |
| Outcome: | The proposed dataset simplifies the development of Named Entity Recognition models for Portuguese literary works. |