Challenge: Large clinical note corpora are one of the most needed and one of least available resources in biomedical NLP due to patient confidentiality considerations and expert annotation cost.
Approach: They present a corpus of 43,985 clinical patient notes (PNs) written by 35,156 examinees during the USMLE® Step 2 Clinical Skills examination.
Outcome: The corpus of 43,985 clinical patient notes (PNs) written by 35,156 examinees during the high-stakes USMLE® Step 2 Clinical Skills examination is available via a data sharing agreement with NBME .

Similar Papers

The Medical Scribe: Corpus Development and Model Performance Analyses (2020.lrec-1)

Copied to clipboard

Challenge: Existing tools to assist in clinical note generation using audio of provider-patient encounters are lacking.
Approach: They develop an annotation scheme to extract relevant clinical concepts from audio of provider-patient encounters and train a state-of-the-art tagging model.
Outcome: The proposed model is more useful than the F-scores reflect and can be used in clinical notes.
A Corpus for Detecting High-Context Medical Conditions in Intensive Care Patient Notes Focusing on Frequently Readmitted Patients (2020.lrec-1)

Copied to clipboard

Challenge: Currently, most medical data is generated and stored in unstructured, text-based format.
Approach: They propose to use a patient phenotyping dataset to identify whether a given medical condition is present in their notes.
Outcome: The proposed dataset contains 1102 Discharge Summaries and 1000 Nursing Progress Notes.
Automatic rubric-based content grading for clinical notes (D19-62)

Copied to clipboard

Challenge: Clinical notes are important documentation critical to medical care, as well as billing and legal needs.
Approach: They propose to evaluate clinical note creation using rubric-based content grading . they build a feature-based system and a neural network-based baseline system .
Outcome: The proposed system can be used to evaluate clinical notes for a rubric-based content grading system . the proposed system has content point accuracy and kappa values at 0.86 and 0.71 on the test set .
MIMICause: Representation and automatic extraction of causal relation types from clinical notes (2022.findings-acl)

Copied to clipboard

Challenge: Extracted causal information from clinical notes can be combined with structured EHR data such as demographics, diagnoses, and medications.
Approach: They propose to annotate clinical notes and develop an annotated corpus and provide baseline scores to identify types and direction of causal relations between a pair of biomedical concepts.
Outcome: The proposed annotation guidelines achieved a high inter-annotator agreement and a macro F1 score on the clinical text.
Annotation of a Large Clinical Entity Corpus (D18-1)

Copied to clipboard

Challenge: Past researches have shown the superiority of statistical/ML approaches over the rule based approaches.
Approach: They propose to annotate a clinical domain annotated corpus using a small data set or a narrower domain to take full advantage of machine learning.
Outcome: The proposed corpus contains 5,160 clinical documents from forty different clinical specialties.
An Empirical Study of Clinical Note Generation from Doctor-Patient Encounters (2023.eacl-main)

Copied to clipboard

Challenge: Medical doctors spend 52 to 102 minutes per day writing clinical notes from patient encounters.
Approach: They propose to use a new dataset to generate automated and manual clinical notes from doctor-patient conversations in a clinical setting.
Outcome: The proposed model could reduce the time spent writing clinical notes from doctor-patient conversations in a clinical setting.
emrQA: A Large Corpus for Question Answering on Electronic Medical Records (D18-1)

Copied to clipboard

Challenge: Existing annotations for other NLP tasks are used to generate domain-specific large-scale question answering (QA) datasets.
Approach: They propose to re-purpose existing annotations for other NLP tasks by generating a large-scale question answering corpus using 1 million questions-logical form and 400,000+ question-answer evidence pairs.
Outcome: The proposed model can be trained to learn domain-specific large-scale question answering (QA) datasets.
A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature (P18-1)

Copied to clipboard

Challenge: In 2015 alone, about 100 manuscripts describing randomized controlled trials for medical interventions were published every day.
Approach: They propose a corpus of 5,000 medical articles annotated with demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured.
Outcome: The proposed corpus includes 5,000 medical articles describing clinical randomized controlled trials.
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes (2024.findings-acl)

Copied to clipboard

Challenge: Clinical notes are an extensive repository of information specific to individual patients.
Approach: They create synthetic large-scale clinical notes using publicly available case reports extracted from biomedical literature and train a clinical large language model, Asclepius.
Outcome: The proposed model outperforms several other models and is supported by detailed evaluations conducted by GPT-4 and medical professionals.
An Investigation of Evaluation Methods in Automatic Medical Note Generation (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies show that doctors can save significant amounts of time when using automatic note generation.
Approach: They propose task-specific metrics for automatic note generation from medical conversation summarization and generation, including knowledge-graph embedding-based metrics, customized model-based measures with domain-specific weights, and ensemble metrics.
Outcome: The proposed evaluation metrics are compared to existing models and can have different behaviors on different types of clinical notes datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations