Challenge: Biomedical natural language processing (BioNLP) has long been recognized as effective method to accelerate drug-related knowledge discovery.
Approach: They developed an active gene annotation corpus (AGAC) to support drug repurposing.
Outcome: The active gene annotation corpus (AGAC) was developed to support knowledge discovery for drug repurposing.

Similar Papers

Biomedical relation extraction with pre-trained language representations and minimal task-specific architecture (D19-57)

Copied to clipboard

Challenge: Using a pre-trained BERT-Base model, we learn domain-specific language representations using biomedical text.
Approach: They propose a system that extends BERT, a state-of-the-art language model, which learns contextual language representations from a large unlabelled corpus.
Outcome: The proposed model outperforms a baseline model while relying on an extremely simple setup with no specially engineered features.
Trigger Word Detection and Thematic Role Identification via BERT and Multitask Learning (D19-57)

Copied to clipboard

Challenge: Using natural language processing to discover and mine drug-related knowledge from text has been a hot topic in recent years.
Approach: They propose to use a pre-trained biomedical language representation model to extract mutation-disease knowledge from PubMed.
Outcome: The proposed approaches achieve 0.60 (ranks 1) and 0.25 (rank 2) on task 1 and task 2 respectively in terms of F1 metric.
DeepGeneMD: A Joint Deep Learning Model for Extracting Gene Mutation-Disease Knowledge from PubMed Literature (D19-57)

Copied to clipboard

Challenge: Identifying and understanding the pathogenesis of genetic diseases is an essential task.
Approach: They propose a joint deep learning model for gene mutation-disease knowledge extraction that adapts the state-of-the-art hierarchical multi-task learning framework for joint inference on named entity recognition and relation extraction.
Outcome: The proposed model achieves the average score of 0.45 on recognizing gene activities and disease entities and the average F1 score of 0.3 on extracting relations, ranking 1st in the AGAC RE task.
BioT5+: Towards Generalized Biological Understanding with IUPAC Integration and Multi-task Tuning (2024.findings-acl)

Copied to clipboard

Challenge: BioT5+ is an extension of the BioT5, but lacked a nuanced understanding of molecular structures.
Approach: They propose a new bio-entity modeling framework, BioT5+, which integrates IUPAC names and molecule data.
Outcome: The proposed model bridges the gap between molecular representations and textual descriptions and improves the grounded reasoning of bio-text and bio-sequences.
Infrastructure for Semantic Annotation in the Genomics Domain (2020.lrec-1)

Copied to clipboard

Challenge: a novel infrastructure for biomedical text mining combines NLP and corpus linguistics methods to provide a comprehensive corpus for literature-based discovery.
Approach: They propose a novel pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature . it uses an updatable Gene Ontology Semantic Tagger and a NLP pipeline scheduler to collect and process the corpus.
Outcome: The proposed infrastructure allows for extreme-scale research on the open access PubMed Central archive.
ProGene - A Large-scale, High-Quality Protein-Gene Annotated Benchmark Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Genes and proteins are fundamental entities of molecular genetics and are important for precision medicine.
Approach: They propose to use a corpus of gene and protein names to cope with this class of named entities in a large-scale annotation campaign at the Jena University Language & Information Engineering lab.
Outcome: The proposed corpus is an overall subdomain-independent corpus . it consists of 3,308 MEDLINE abstracts with over 36k sentences and more than 960k tokens annotated with nearly 60k named entity mentions.
A New Corpus to Support Text Mining for the Curation of Metabolites in the ChEBI Database (L18-1)

Copied to clipboard

Challenge: a corpus of 200 abstracts and 100 full text papers which have been annotated with named entities and relations in the biomedical domain is part of the OpenMinTeD project.
Approach: They propose to annotate 200 abstracts and 100 full text papers with entities and relations in the biomedical domain as part of the OpenMinTeD project.
Outcome: The proposed corpus can be used within ChEBI to facilitate text and data mining and integrate with the OpenMinTeD text and database platform.
RDoC Task at BioNLP-OST 2019 (D19-57)

Copied to clipboard

Challenge: BioNLP-OST is an international competition organized to facilitate development and sharing of computational tasks of biomedical text mining and solutions to them.
Approach: They propose a new mental health informatics task that is composed of two subtasks: information retrieval and sentence extraction.
Outcome: The proposed task performed well on both tasks, but there are still challenges.
Adverse Event Extraction from Discharge Summaries: A New Dataset, Annotation Scheme, and Initial Findings (2025.acl-long)

Copied to clipboard

Challenge: Existing resources for AE extraction are limited due to complexity, variability, and ambiguity of clinical narratives.
Approach: They present a manually annotated corpus for Adverse Event (AE) extraction from discharge summaries of elderly patients.
Outcome: The proposed model performs well on coarse-grained extraction, but drops notably for rare events and complex attributes.
Proceedings of the 5th Workshop on BioNLP Open Shared Tasks (D19-57)

Copied to clipboard

Challenge: a workshop organized by BioNLP-ST aims to share computational tasks of biomedical text mining and solutions to them.
Approach: this year, six tasks are contributed by voluntary task organizers . they aim to promote the sharing of computational tasks of biomedical text mining . 43 reviewers selected 30 papers to be presented for the workshop .
Outcome: the BioNLP Open Shared Tasks is organized to promote the sharing of computational tasks of biomedical text mining and solutions to them.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations