Challenge: Existing medical social media corpora focus on a small set of entities and relations . existing text mining and information extraction methods focus on scientific text generated by researchers but their access to individual patient experiences or patient-doctor interactions is limited.
Approach: The dataset consists of 2,100 medical tweets with approx. 6,000 entities and 2,200 relations.
Outcome: The proposed dataset consists of 2,100 tweets with approx. 6,000 entities and 2,200 relations.

Similar Papers

COMETA: A Corpus for Medical Entity Linking in the Social Media (2020.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for Entity Linking (EL) fail to address the complex nature of health terminology in layman’s language.
Approach: They propose to use a corpus of 20k English biomedical entity mentions from Reddit expert-annotated with links to a widely-used medical knowledge graph to investigate the ability of these systems to perform complex inference on entities and concepts.
Outcome: The proposed corpus satisfies a combination of desirable properties, from scale and coverage to diversity and quality, that to the best of our knowledge has not been met by existing resources in the field.
Named Entities in Medical Case Reports: Corpus and Experiments (2020.lrec-1)

Copied to clipboard

Challenge: Only very few annotated corpora in the medical domain exist.
Approach: They propose to annotate medical entities in case reports from PubMed Central's open access library.
Outcome: The proposed corpus is the first of its kind to be made available to the scientific community in English.
Learning to Leverage High-Order Medical Knowledge Graph for Joint Entity and Relation Extraction (2023.findings-acl)

Copied to clipboard

Challenge: Medical terms are difficult to understand and relations between medical entities become complicated.
Approach: They propose to leverage medical domain knowledge for extracting entities and relations for Chinese medical texts by building a heterogeneous graph based on medical knowledge graph.
Outcome: The proposed method is more effective than state-of-the-art methods on real Chinese medical texts.
RedHOT: A Corpus of Annotated Medical Questions, Experiences, and Claims on Social Media (2023.findings-eacl)

Copied to clipboard

Challenge: Social media platforms such as Reddit are vulnerable to misinformation and disinformation.
Approach: They propose a method to automatically derive (noisy) supervision for retrieval of trustworthy evidence relevant to a given claim made on social media.
Outcome: The proposed method outperforms baseline models in the retrieval task performed by medical doctors.
Relating Relations: Meta-Relation Extraction from Online Health Forum Posts (2021.eacl-srw)

Copied to clipboard

Challenge: Relation extraction is a key task in knowledge extraction, and is often defined as identifying relations that hold between entities in text.
Approach: They propose to conceptualise relation extraction tasks for user-generated health texts and create a dataset and model for meta-relation extraction.
Outcome: The proposed model will be able to extract meta-relations from user-generated health texts with tolerable cognitive load and a new dataset and annotation scheme with tolerance for annotations.
Towards Extracting Medical Family History from Natural Language Interactions: A New Dataset and Baselines (D19-1)

Copied to clipboard

Challenge: Using dialog agents, we can collect family history data from in-person consultations and crowdsource it to a genetic counselor.
Approach: They propose to use natural language interactions annotated with medical family histories to collect information from a genetic counselor and crowdsourcing.
Outcome: The proposed system averages 0.87 on complex sentences on the targeted relations.
Method Entity Extraction from Biomedical Texts (2022.coling-1)

Copied to clipboard

Challenge: Scientific research papers consist of complex keywords and domain-specific terminologies, and new terminologie erupt.
Approach: They find method terminologies in biomedical text using rule-based and machine learning techniques . authors propose to use a silver standard corpus to extract method entities from biomedically text .
Outcome: The proposed method entities can be extracted from biomedical text with reasonable accuracy . the proposed method entity extraction method is based on a rule-based method and a machine learning technique.
Medical Sentiment Analysis using Social Media: Towards building a Patient Assisted System (L18-1)

Copied to clipboard

Challenge: a study conducted by the pew Internet & American Life Project 1 shows that almost 80 percent of Internet users have explored health-related topic online.
Approach: They propose to crawl medical forums with opinions about medical condition self narrated by users.
Outcome: The proposed system is based on opinions about medical condition self-narrated by users on medical forums.
LENS: Learning Entities from Narratives of Skin Cancer (2025.coling-demos)

Copied to clipboard

Challenge: Learning entities from narratives of skin cancer (LENS) is an automatic entity recognition system built on colloquial writings from skin cancer-related forums.
Approach: They propose to use reddit forums to create an automatic entity recognition system that can be used to predict skin cancer outcomes.
Outcome: LENS achieves an overall entity-level F1 score of 0.561 . other notable results include “CANC_T” (0.747), “STG” (0.888), “POB” (0.914), “GENDER” (0.750), “A/G” (00.646), “EMO” (0.619), and “MHD” (0.503).
Accelerating the Discovery of Semantic Associations from Medical Literature: Mining Relations Between Diseases and Symptoms (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing methods to extract semantic associations from medical literature do not take into account the semantics of sentences from which entity co-occurrences are extracted.
Approach: They propose a system for the automatic discovery of semantic associations between different entities such as diseases and their symptoms using a semantic network and a binary relation classification model trained with distant supervision.
Outcome: The proposed system validates the extracted associations against a publicly available list of disease-symptom pairs against 14M PubMed abstracts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations