Profiling Medical Journal Articles Using a Gene Ontology Semantic Tagger (L18-1)
Copied to clipboard
| Challenge: | a growing number of scientific publications are based on sub-divisions and sub-communities of expertise becoming disconnected from each other. |
| Approach: | They propose to examine corpora derived from bodies of genetics literature and use it to make comparisons and improve retrieval methods. |
| Outcome: | The proposed methods will help to make comparisons and improve retrieval methods using domain knowledge via an existing gene ontology. |
Similar Papers
Infrastructure for Semantic Annotation in the Genomics Domain (2020.lrec-1)
Copied to clipboard
Mahmoud El-Haj, Nathan Rutherford, Matthew Coole, Ignatius Ezeani, Sheryl Prentice, Nancy Ide, Jo Knight, Scott Piao, John Mariani, Paul Rayson, Keith Suderman
| Challenge: | a novel infrastructure for biomedical text mining combines NLP and corpus linguistics methods to provide a comprehensive corpus for literature-based discovery. |
| Approach: | They propose a novel pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature . it uses an updatable Gene Ontology Semantic Tagger and a NLP pipeline scheduler to collect and process the corpus. |
| Outcome: | The proposed infrastructure allows for extreme-scale research on the open access PubMed Central archive. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
Automatic Term Name Generation for Gene Ontology: Task and Dataset (2020.findings-emnlp)
Copied to clipboard
Yanjian Zhang, Qin Chen, Yiteng Zhang, Zhongyu Wei, Yixu Gao, Jiajie Peng, Zengfeng Huang, Weijian Sun, Xuanjing Huang
| Challenge: | Gene Ontology (GO) terms are used to describe gene function in biology and bio-medicine. |
| Approach: | They propose a task to generate term names for GO and build a large-scale benchmark dataset. |
| Outcome: | The proposed model outperforms baselines by incorporating the relations between genes, words and terms for term name generation. |
Accelerating the Discovery of Semantic Associations from Medical Literature: Mining Relations Between Diseases and Symptoms (2022.emnlp-industry)
Copied to clipboard
| Challenge: | Existing methods to extract semantic associations from medical literature do not take into account the semantics of sentences from which entity co-occurrences are extracted. |
| Approach: | They propose a system for the automatic discovery of semantic associations between different entities such as diseases and their symptoms using a semantic network and a binary relation classification model trained with distant supervision. |
| Outcome: | The proposed system validates the extracted associations against a publicly available list of disease-symptom pairs against 14M PubMed abstracts. |
Named Entities in Medical Case Reports: Corpus and Experiments (2020.lrec-1)
Copied to clipboard
| Challenge: | Only very few annotated corpora in the medical domain exist. |
| Approach: | They propose to annotate medical entities in case reports from PubMed Central's open access library. |
| Outcome: | The proposed corpus is the first of its kind to be made available to the scientific community in English. |
ProGene - A Large-scale, High-Quality Protein-Gene Annotated Benchmark Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Genes and proteins are fundamental entities of molecular genetics and are important for precision medicine. |
| Approach: | They propose to use a corpus of gene and protein names to cope with this class of named entities in a large-scale annotation campaign at the Jena University Language & Information Engineering lab. |
| Outcome: | The proposed corpus is an overall subdomain-independent corpus . it consists of 3,308 MEDLINE abstracts with over 36k sentences and more than 960k tokens annotated with nearly 60k named entity mentions. |
Structured Multi-Label Biomedical Text Tagging via Attentive Neural Tree Decoding (D18-1)
Copied to clipboard
| Challenge: | Existing methods for tagging unstructured texts with arbitrary number of terms drawn from an ontology are lacking. |
| Approach: | They propose a model for tagging unstructured texts with an arbitrary number of terms drawn from an ontology. |
| Outcome: | The proposed model yields state-of-the-art results on the important task of assigning MeSH terms to biomedical abstracts. |
Fine-grained Information Extraction from Biomedical Literature based on Knowledge-enriched Abstract Meaning Representation (2021.acl-long)
Copied to clipboard
| Challenge: | Compared with general natural language texts, sentences from scientific papers usually possess wider contexts between knowledge elements. |
| Approach: | They propose a novel biomedical Information Extraction model to extract scientific entities and events from English research papers using Abstract Meaning Representation (AMR) they construct a sentence-level knowledge graph from an external knowledge base and encode it to improve the model's understanding of complex scientific concepts. |
| Outcome: | The proposed model can extract scientific entities and events from scientific literature and improve its understanding of complex scientific concepts. |
A Silver Standard Corpus of Human Phenotype-Gene Relations (N19-1)
Copied to clipboard
| Challenge: | Existing tools for phenotype-gene relations extraction require annotated corpus, which requires manual effort and time. |
| Approach: | They propose to generate a silver standard corpus of human phenotype and gene annotations and their relations using Named-Entity Recognition tools. |
| Outcome: | The proposed corpus was generated with Named-Entity Recognition tools with a precision of 87.01%. |
The Medical Scribe: Corpus Development and Model Performance Analyses (2020.lrec-1)
Copied to clipboard
Izhak Shafran, Nan Du, Linh Tran, Amanda Perry, Lauren Keyes, Mark Knichel, Ashley Domin, Lei Huang, Yu-hui Chen, Gang Li, Mingqiu Wang, Laurent El Shafey, Hagen Soltau, Justin Stuart Paul
| Challenge: | Existing tools to assist in clinical note generation using audio of provider-patient encounters are lacking. |
| Approach: | They develop an annotation scheme to extract relevant clinical concepts from audio of provider-patient encounters and train a state-of-the-art tagging model. |
| Outcome: | The proposed model is more useful than the F-scores reflect and can be used in clinical notes. |