Challenge: a low overall error rate is unavoidable in the field of bioNLP, argues a new study.
Approach: They propose a highly extensible open source framework that supports BioNLP for the pharmaceutical sector.
Outcome: The proposed framework is extensible and scalable and can support BioNLP for the pharmaceutical sector.

Similar Papers

BioEL: A Comprehensive Python Package for Biomedical Entity Linking (2025.findings-naacl)

Copied to clipboard

Challenge: Entity Linking in biomedical literature is a critical task that enhances the extraction and integration of information from diverse scientific literature.
Approach: They propose a Python package that allows for better Entity Linking in biomedical literature . the package includes four components: Ontology Object, Dataset Object and Evaluation Framework .
Outcome: The proposed open-source package enables the implementation and comparison of biomedical entity linking tasks.
MedCATTrainer: A Biomedical Free Text Annotation Interface with Active Learning and Research Use Case Specific Customisation (D19-3)

Copied to clipboard

Challenge: 80% of biomedical data is stored in unstructured text such as electronic health records (EHRs).
Approach: They propose a web-based interface for building, improving and customising a given Named Entity Recognition and Linking (NER+L) model for biomedical domain text.
Outcome: The proposed interface is designed to build, improve and customise a NER+L model for biomedical domain text and collate accurate research use case specific training data.
INSIGHTBUDDY-AI: Medication Extraction and Entity Linking using Pre-Trained Language Models and Ensemble Learning (2025.naacl-srw)

Copied to clipboard

Challenge: InsightBuddy-AI is a system for extracting medication mentions and their associated attributes.
Approach: They propose a system for extracting medication mentions and their associated attributes . they use stacked and voting ensembles built upon pre-trained language models .
Outcome: The proposed system outperforms fine-tuned models in the extraction of medication mentions and associated attributes.
Named Entity Recognition for Chinese biomedical patents (2020.coling-main)

Copied to clipboard

Challenge: Existing attempts to address NER for Chinese biomedical texts have been limited due to the amount of Chinese biomedicine discoveries being patented.
Approach: They train and evaluate Chinese biomedical patents NER models based on BERT . their model is optimized for Chinese bio-patent data and scored an F1 .
Outcome: The proposed model achieves an F1 score of 0.540.15 for Chinese biomedical patent data.
GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers (2022.lrec-1)

Copied to clipboard

Challenge: despite advances in language resources, there is still a shortage of annotated corpora covering (German) medical language.
Approach: They propose to build on clinical guidelines with an annotation scheme based on SNOMED CT . they also train named entity recognition models on the new data set .
Outcome: The new corpus can be built upon clinical guidelines with reasonable coverage of medical terminology.
BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for multi-hop reasoning in biomedical domain are lacking . bioHopR provides benchmarks to evaluate multi-step reasoning in structured biomedic knowledge graphs .
Approach: They propose a benchmark to evaluate multi-hop, multi-answer reasoning in biomedical knowledge graphs.
Outcome: BioHopR evaluates multi-hop reasoning in biomedical knowledge graphs based on the PrimeKG model . it outperforms proprietary models and open-source biomedal models in 1-hop and 2-hop tasks .
A Deep Learning-Based System for PharmaCoNER (D19-57)

Copied to clipboard

Challenge: Efficient access to mentions of clinical entities is very important for using clinical text.
Approach: They developed a pipeline system based on deep learning methods for this shared task . it achieves a micro-average F1-score of 0.9105 on track 1 and a mini-average LSTM score of 0.8391 on track 2 .
Outcome: The proposed system achieves a micro-average F1-score of 0.9105 on track 1 and a mini-average score of 0.8391 on track 2.
OpenBioNER: Lightweight Open-Domain Biomedical Named Entity Recognition Through Entity Type Description (2025.findings-naacl)

Copied to clipboard

Challenge: Biomedical Named Entity Recognition (BioNER) is a computationally expensive and limited tool . specialized 7B NER LLMs and GPT-4o can't match textual spans with entity types .
Approach: They propose a lightweight BERT-based cross-encoder architecture that can identify any biomedical entity using only its description.
Outcome: The proposed system outperforms existing models that match textual spans with entity types rather than descriptions on biomedical benchmarks.
An End-to-End Progressive Multi-Task Learning Framework for Medical Named Entity Recognition and Normalization (2021.acl-long)

Copied to clipboard

Challenge: Existing models for medical named entity recognition and named entity normalization suffer from error propagation between the two tasks.
Approach: They propose an end-to-end progressive multi-task learning model for jointly modeling medical named entity recognition and normalization in an effective way.
Outcome: The proposed model reduces error propagation by exploiting the learnable features for both tasks to improve performance.
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment (2025.acl-long)

Copied to clipboard

Challenge: Large language models such as GPT-4 have limited their deployment in clinical settings . a novel framework for adapting SLMs into high-performing clinical models is needed .
Approach: They propose a framework for adapting large language models into high-performing clinical models . they pre-instruct experts on relevant medical and clinical corpora and model merging .
Outcome: The proposed framework outperforms the existing model on the CLUE+ benchmark on medical entities and radiology reports.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations