Automatic rubric-based content grading for clinical notes (D19-62)

Copied to clipboard

Challenge: Clinical notes are important documentation critical to medical care, as well as billing and legal needs.
Approach: They propose to evaluate clinical note creation using rubric-based content grading . they build a feature-based system and a neural network-based baseline system .
Outcome: The proposed system can be used to evaluate clinical notes for a rubric-based content grading system . the proposed system has content point accuracy and kappa values at 0.86 and 0.71 on the test set .

Similar Papers

An Investigation of Evaluation Methods in Automatic Medical Note Generation (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies show that doctors can save significant amounts of time when using automatic note generation.
Approach: They propose task-specific metrics for automatic note generation from medical conversation summarization and generation, including knowledge-graph embedding-based metrics, customized model-based measures with domain-specific weights, and ensemble metrics.
Outcome: The proposed evaluation metrics are compared to existing models and can have different behaviors on different types of clinical notes datasets.
The Medical Scribe: Corpus Development and Model Performance Analyses (2020.lrec-1)

Copied to clipboard

Challenge: Existing tools to assist in clinical note generation using audio of provider-patient encounters are lacking.
Approach: They develop an annotation scheme to extract relevant clinical concepts from audio of provider-patient encounters and train a state-of-the-art tagging model.
Outcome: The proposed model is more useful than the F-scores reflect and can be used in clinical notes.
Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation (2022.acl-long)

Copied to clipboard

Challenge: Recent studies suggest that note generation systems can be used to generate clinical consultation notes from the verbatim transcript of the consultation.
Approach: They propose to use machine learning to generate consultation notes from the verbatim transcript of the consultation to evaluate their effectiveness.
Outcome: The proposed model performs better than common model-based metrics like BertScore and is open-sourced.
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing automated metrics fail to align with real-world physician preferences.
Approach: They propose a pipeline that distills real user feedback into structured checklists for note evaluation that are interpretable, grounded in human feedback, and enforceable by LLM-based evaluators.
Outcome: The proposed checklist outperforms baseline evaluations in coverage, diversity, and predictive power for human ratings.
The USMLE® Step 2 Clinical Skills Patient Note Corpus (2022.naacl-main)

Copied to clipboard

Challenge: Large clinical note corpora are one of the most needed and one of least available resources in biomedical NLP due to patient confidentiality considerations and expert annotation cost.
Approach: They present a corpus of 43,985 clinical patient notes (PNs) written by 35,156 examinees during the USMLE® Step 2 Clinical Skills examination.
Outcome: The corpus of 43,985 clinical patient notes (PNs) written by 35,156 examinees during the high-stakes USMLE® Step 2 Clinical Skills examination is available via a data sharing agreement with NBME .
Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing studies have shown that note generation is difficult due to subjective nature of many aspects of output quality.
Approach: They propose a protocol that aims to increase objectivity by grounding evaluations in Consultation Checklists, which are created in a preliminary step and then used as a common point of reference during quality assessment.
Outcome: The proposed protocol shows that the evaluations produced in the study are more objective than the original human note.
Extracting relevant information from physician-patient dialogues for automated clinical note taking (D19-62)

Copied to clipboard

Challenge: a system that extracts pertinent medical information from dialogues between clinicians and patients is proposed . entering data into EMRs is currently slow and error-prone, and clinicians spend up to 50% of their time on data entry.
Approach: They propose a system that automatically extracts medical information from dialogues between clinicians and patients using context and time information.
Outcome: The proposed system extracts medical information from dialogues and automatically generates a patient note.
TN-Eval: Rubric and Evaluation Protocols for Measuring the Quality of Behavioral Therapy Notes (2025.acl-industry)

Copied to clipboard

Challenge: Behavioral therapy notes are important for legal compliance and patient care, but quality standards for them remain underdeveloped.
Approach: They propose a rubric for evaluating therapy notes across key dimensions: completeness, conciseness, faithfulness.
Outcome: The proposed evaluation framework improves on therapist-written notes and LLM-generated notes.
ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical Judgment (2025.emnlp-main)

Copied to clipboard

Challenge: Automatically generated radiology reports often receive high scores from existing evaluation metrics but fail to earn clinicians’ trust.
Approach: They propose a meta-evaluation framework that uses criteria spanning discrimination, robustness, and monotonicity to evaluate existing metrics.
Outcome: The proposed framework offers guidance for building more clinically reliable evaluation methods.
Generating Accurate Electronic Health Assessment from Medical Graph (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models based on medical domain-specific knowledge or patients’ prior diagnoses and clinical encounters were mainly based upon clinical diagnoses.
Approach: They propose a graph neural network model that incorporates clinical knowledge into an end-to-end corpus-learning system and builds on it.
Outcome: The proposed model significantly improves the BLEU and rouge score compared with baseline models and physicians’ evaluation showed that it generates high-quality assessments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations