A New Corpus to Support Text Mining for the Curation of Metabolites in the ChEBI Database (L18-1)
Copied to clipboard
Matthew Shardlow, Nhung Nguyen, Gareth Owen, Claire O’Donovan, Andrew Leach, John McNaught, Steve Turner, Sophia Ananiadou
| Challenge: | a corpus of 200 abstracts and 100 full text papers which have been annotated with named entities and relations in the biomedical domain is part of the OpenMinTeD project. |
| Approach: | They propose to annotate 200 abstracts and 100 full text papers with entities and relations in the biomedical domain as part of the OpenMinTeD project. |
| Outcome: | The proposed corpus can be used within ChEBI to facilitate text and data mining and integrate with the OpenMinTeD text and database platform. |
Similar Papers
Named Entities in Medical Case Reports: Corpus and Experiments (2020.lrec-1)
Copied to clipboard
| Challenge: | Only very few annotated corpora in the medical domain exist. |
| Approach: | They propose to annotate medical entities in case reports from PubMed Central's open access library. |
| Outcome: | The proposed corpus is the first of its kind to be made available to the scientific community in English. |
Infrastructure for Semantic Annotation in the Genomics Domain (2020.lrec-1)
Copied to clipboard
Mahmoud El-Haj, Nathan Rutherford, Matthew Coole, Ignatius Ezeani, Sheryl Prentice, Nancy Ide, Jo Knight, Scott Piao, John Mariani, Paul Rayson, Keith Suderman
| Challenge: | a novel infrastructure for biomedical text mining combines NLP and corpus linguistics methods to provide a comprehensive corpus for literature-based discovery. |
| Approach: | They propose a novel pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature . it uses an updatable Gene Ontology Semantic Tagger and a NLP pipeline scheduler to collect and process the corpus. |
| Outcome: | The proposed infrastructure allows for extreme-scale research on the open access PubMed Central archive. |
PharmaCoNER: Pharmacological Substances, Compounds and proteins Named Entity Recognition track (D19-57)
Copied to clipboard
Aitor Gonzalez-Agirre, Montserrat Marimon, Ander Intxaurrondo, Obdulia Rabal, Marta Villegas, Martin Krallinger
| Challenge: | Biomedical text mining is one of the most prolific application domains of natural language processing technologies. |
| Approach: | They propose to share a task on detecting drug and chemical entities in medical documents in Spanish with other languages to improve access to biomedical text mining. |
| Outcome: | The first task on detecting drug and chemical entities in Spanish medical documents yielded competitive results with F-measures above 0.91. |
A Distant Supervision Corpus for Extracting Biomedical Relationships Between Chemicals, Diseases and Genes (2022.lrec-1)
Copied to clipboard
| Challenge: | Biomedical researchers have used manual curation to extract biomedical interactions from research texts to improve coverage. |
| Approach: | They propose a new dataset for training and evaluating multi-class multi-label biomedical relation extraction models using human annotations and the CTD database. |
| Outcome: | The proposed dataset is substantially larger and cleaner than existing datasets and includes annotations linking mentions to their entities. |
S2ORC: The Semantic Scholar Open Research Corpus (2020.acl-main)
Copied to clipboard
| Challenge: | Academic papers are an increasingly important textual domain for natural language processing (NLP) research. |
| Approach: | They propose to aggregate 81.1M English-language academic papers into a unified source . they hope this resource will facilitate research and development of tools for text mining over academic text. |
| Outcome: | The proposed corpus includes metadata, abstracts, bibliographic references, and structured full text for 8.1M open access papers. |
ChEMU-Ref: A Corpus for Modeling Anaphora Resolution in the Chemical Domain (2021.eacl-main)
Copied to clipboard
| Challenge: | Using a novel annotation scheme, we identify anaphoric references in chemical patents and determine the chemical relation between linked entities. |
| Approach: | They propose a neural approach to anaphora resolution based on coreference and bridging links in chemical patents. |
| Outcome: | The proposed framework can be used to identify anaphoric references in chemical patents and determine the chemical relation between linked entities. |
Text Mining for History: first steps on building a large dataset (L18-1)
Copied to clipboard
| Challenge: | a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way . |
| Approach: | They propose to use a Brazilian historical-biographical dictionary as a resource for text mining. |
| Outcome: | The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated . |
Annotation of a Large Clinical Entity Corpus (D18-1)
Copied to clipboard
| Challenge: | Past researches have shown the superiority of statistical/ML approaches over the rule based approaches. |
| Approach: | They propose to annotate a clinical domain annotated corpus using a small data set or a narrower domain to take full advantage of machine learning. |
| Outcome: | The proposed corpus contains 5,160 clinical documents from forty different clinical specialties. |
Recovering Patient Journeys: A Corpus of Biomedical Entities and Relations on Twitter (BEAR) (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing medical social media corpora focus on a small set of entities and relations . existing text mining and information extraction methods focus on scientific text generated by researchers but their access to individual patient experiences or patient-doctor interactions is limited. |
| Approach: | The dataset consists of 2,100 medical tweets with approx. 6,000 entities and 2,200 relations. |
| Outcome: | The proposed dataset consists of 2,100 tweets with approx. 6,000 entities and 2,200 relations. |
BioRo: The Biomedical Corpus for the Romanian Language (L18-1)
Copied to clipboard
| Challenge: | Biomedical text mining uses linguistic resources available in English, but for other languages such as Romanian, the access to language resources is not straight-forward. |
| Approach: | They present a biomedical corpus of the Romanian language, which is a valuable linguistic asset for biomedically text mining. |
| Outcome: | The proposed corpus will be made publicly available to the biomedical text mining community . the corpus is a reference corpus for the Romanian language . |