Challenge: a corpus of 200 abstracts and 100 full text papers which have been annotated with named entities and relations in the biomedical domain is part of the OpenMinTeD project.
Approach: They propose to annotate 200 abstracts and 100 full text papers with entities and relations in the biomedical domain as part of the OpenMinTeD project.
Outcome: The proposed corpus can be used within ChEBI to facilitate text and data mining and integrate with the OpenMinTeD text and database platform.

Similar Papers

Named Entities in Medical Case Reports: Corpus and Experiments (2020.lrec-1)

Copied to clipboard

Challenge: Only very few annotated corpora in the medical domain exist.
Approach: They propose to annotate medical entities in case reports from PubMed Central's open access library.
Outcome: The proposed corpus is the first of its kind to be made available to the scientific community in English.
Infrastructure for Semantic Annotation in the Genomics Domain (2020.lrec-1)

Copied to clipboard

Challenge: a novel infrastructure for biomedical text mining combines NLP and corpus linguistics methods to provide a comprehensive corpus for literature-based discovery.
Approach: They propose a novel pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature . it uses an updatable Gene Ontology Semantic Tagger and a NLP pipeline scheduler to collect and process the corpus.
Outcome: The proposed infrastructure allows for extreme-scale research on the open access PubMed Central archive.
PharmaCoNER: Pharmacological Substances, Compounds and proteins Named Entity Recognition track (D19-57)

Copied to clipboard

Challenge: Biomedical text mining is one of the most prolific application domains of natural language processing technologies.
Approach: They propose to share a task on detecting drug and chemical entities in medical documents in Spanish with other languages to improve access to biomedical text mining.
Outcome: The first task on detecting drug and chemical entities in Spanish medical documents yielded competitive results with F-measures above 0.91.
A Distant Supervision Corpus for Extracting Biomedical Relationships Between Chemicals, Diseases and Genes (2022.lrec-1)

Copied to clipboard

Challenge: Biomedical researchers have used manual curation to extract biomedical interactions from research texts to improve coverage.
Approach: They propose a new dataset for training and evaluating multi-class multi-label biomedical relation extraction models using human annotations and the CTD database.
Outcome: The proposed dataset is substantially larger and cleaner than existing datasets and includes annotations linking mentions to their entities.
S2ORC: The Semantic Scholar Open Research Corpus (2020.acl-main)

Copied to clipboard

Challenge: Academic papers are an increasingly important textual domain for natural language processing (NLP) research.
Approach: They propose to aggregate 81.1M English-language academic papers into a unified source . they hope this resource will facilitate research and development of tools for text mining over academic text.
Outcome: The proposed corpus includes metadata, abstracts, bibliographic references, and structured full text for 8.1M open access papers.
ChEMU-Ref: A Corpus for Modeling Anaphora Resolution in the Chemical Domain (2021.eacl-main)

Copied to clipboard

Challenge: Using a novel annotation scheme, we identify anaphoric references in chemical patents and determine the chemical relation between linked entities.
Approach: They propose a neural approach to anaphora resolution based on coreference and bridging links in chemical patents.
Outcome: The proposed framework can be used to identify anaphoric references in chemical patents and determine the chemical relation between linked entities.
Text Mining for History: first steps on building a large dataset (L18-1)

Copied to clipboard

Challenge: a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way .
Approach: They propose to use a Brazilian historical-biographical dictionary as a resource for text mining.
Outcome: The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated .
Annotation of a Large Clinical Entity Corpus (D18-1)

Copied to clipboard

Challenge: Past researches have shown the superiority of statistical/ML approaches over the rule based approaches.
Approach: They propose to annotate a clinical domain annotated corpus using a small data set or a narrower domain to take full advantage of machine learning.
Outcome: The proposed corpus contains 5,160 clinical documents from forty different clinical specialties.
Recovering Patient Journeys: A Corpus of Biomedical Entities and Relations on Twitter (BEAR) (2022.lrec-1)

Copied to clipboard

Challenge: Existing medical social media corpora focus on a small set of entities and relations . existing text mining and information extraction methods focus on scientific text generated by researchers but their access to individual patient experiences or patient-doctor interactions is limited.
Approach: The dataset consists of 2,100 medical tweets with approx. 6,000 entities and 2,200 relations.
Outcome: The proposed dataset consists of 2,100 tweets with approx. 6,000 entities and 2,200 relations.
BioRo: The Biomedical Corpus for the Romanian Language (L18-1)

Copied to clipboard

Challenge: Biomedical text mining uses linguistic resources available in English, but for other languages such as Romanian, the access to language resources is not straight-forward.
Approach: They present a biomedical corpus of the Romanian language, which is a valuable linguistic asset for biomedically text mining.
Outcome: The proposed corpus will be made publicly available to the biomedical text mining community . the corpus is a reference corpus for the Romanian language .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations