Biomedical Named Entity Recognition with Multilingual BERT (D19-57)

Copied to clipboard

Challenge: a multilingual model is not specifically tailored to either the language nor the application domain.
Approach: They propose a CRF-based baseline approach and multilingual BERT to the task . they achieve an F-score of 88% on the development data and 87% on the test set with BERT .
Outcome: The proposed model achieves an F score of 88% on the development data and 87% on the test set with BERT.

Similar Papers

Transfer Learning in Biomedical Named Entity Recognition: An Evaluation of BERT in the PharmaCoNER task (D19-57)

Copied to clipboard

Challenge: Existing methods for natural language processing are labor-intensive and skill-dependent . Currently, most biomedical natural language tasks focus on English documents .
Approach: They introduce a BERT benchmark to facilitate the research of PharmaCoNER task . they evaluate two baselines based on Multilingual BERT and BioBERT on the corpus .
Outcome: The proposed task is based on multilingual BERT and BioBERT on the PharmaCoNER corpus.
When Specialization Helps: Using Pooled Contextualized Embeddings to Detect Chemical and Biomedical Entities in Spanish (D19-57)

Copied to clipboard

Challenge: Existing work on pharmacological entities requires manual annotation of these units.
Approach: They propose an approach to task 1 of the PharmaCoNER Challenge to recognize pharmacological entities on a spanish corpus.
Outcome: The proposed approach achieves 89.76% score on a spanish corpus based on pre-trained embeddings and 90.52% score on domain-specific embeddables.
OpenBioNER: Lightweight Open-Domain Biomedical Named Entity Recognition Through Entity Type Description (2025.findings-naacl)

Copied to clipboard

Challenge: Biomedical Named Entity Recognition (BioNER) is a computationally expensive and limited tool . specialized 7B NER LLMs and GPT-4o can't match textual spans with entity types .
Approach: They propose a lightweight BERT-based cross-encoder architecture that can identify any biomedical entity using only its description.
Outcome: The proposed system outperforms existing models that match textual spans with entity types rather than descriptions on biomedical benchmarks.
Where do LLMs currently stand on biomedical NER in both clean and noisy settings ? (2026.findings-eacl)

Copied to clipboard

Challenge: despite advances in medicine, many diseases remain without effective treatments . clinical meta-analysis is essential for drug discovery and clinical research .
Approach: They investigate the performance of large language models (LLMs) on biomedical NER tasks . findings suggest LLMs exhibit a notable degree of robustness to noise .
Outcome: The proposed models are closing the performance gap with BERT-based models and demonstrate particular strengths in low-data settings.
A Benchmark Evaluation of Clinical Named Entity Recognition in French (2024.lrec-main)

Copied to clipboard

Challenge: Masked Language Models (MLMs) have shown strong performance on many NLP tasks.
Approach: They evaluate masked language models for biomedical French on the task of clinical named entity recognition using gold-standard corpora.
Outcome: The proposed model outperforms standard models on the task of clinical named entity recognition in biomedical French while remaining lighter than current models.
TeluguNER: Leveraging Multi-Domain Named Entity Recognition with Deep Transformers (2022.acl-srw)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a successful and well-researched problem in English due to the availability of resources.
Approach: They propose to use two annotated NER datasets for the Telugu language . they compare the finetuned Telugus model with the existing model in NER .
Outcome: The proposed models outperform existing models on a large dataset of 38,363 sentences on telugu and other languages.
Named Entity Recognition for Chinese biomedical patents (2020.coling-main)

Copied to clipboard

Challenge: Existing attempts to address NER for Chinese biomedical texts have been limited due to the amount of Chinese biomedicine discoveries being patented.
Approach: They train and evaluate Chinese biomedical patents NER models based on BERT . their model is optimized for Chinese bio-patent data and scored an F1 .
Outcome: The proposed model achieves an F1 score of 0.540.15 for Chinese biomedical patent data.
PharmaCoNER: Pharmacological Substances, Compounds and proteins Named Entity Recognition track (D19-57)

Copied to clipboard

Challenge: Biomedical text mining is one of the most prolific application domains of natural language processing technologies.
Approach: They propose to share a task on detecting drug and chemical entities in medical documents in Spanish with other languages to improve access to biomedical text mining.
Outcome: The first task on detecting drug and chemical entities in Spanish medical documents yielded competitive results with F-measures above 0.91.
IxaMed at PharmacoNER Challenge 2019 (D19-57)

Copied to clipboard

Challenge: The aim of this paper is to present our approach in the PharmacoNER 2019 task.
Approach: They propose to use a Bi-LSTM with a CRF to identify named entities from clinical case studies written in Spanish.
Outcome: The proposed approach achieves the best score (86.81 F-Score) combining pretrained word embeddings of Wikipedia and Electronic Health Records with contextual string embedds.
Exploring Cross-sentence Contexts for Named Entity Recognition with BERT (2020.coling-main)

Copied to clipboard

Challenge: Named entity recognition (NER) is often addressed as a sequence classification task with each input consisting of one sentence of text.
Approach: They propose a method to combine different predictions from multiple sentences in input samples to increase NER performance.
Outcome: The proposed method improves on the state-of-the-art NER results on English, Dutch, and Finnish and achieves the best reported BERT-based results on German.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations