Papers by Wen-wai Yim

6 papers
An Investigation of Evaluation Methods in Automatic Medical Note Generation (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies show that doctors can save significant amounts of time when using automatic note generation.
Approach: They propose task-specific metrics for automatic note generation from medical conversation summarization and generation, including knowledge-graph embedding-based metrics, customized model-based measures with domain-specific weights, and ensemble metrics.
Outcome: The proposed evaluation metrics are compared to existing models and can have different behaviors on different types of clinical notes datasets.
Automatic rubric-based content grading for clinical notes (D19-62)

Copied to clipboard

Challenge: Clinical notes are important documentation critical to medical care, as well as billing and legal needs.
Approach: They propose to evaluate clinical note creation using rubric-based content grading . they build a feature-based system and a neural network-based baseline system .
Outcome: The proposed system can be used to evaluate clinical notes for a rubric-based content grading system . the proposed system has content point accuracy and kappa values at 0.86 and 0.71 on the test set .
An Empirical Study of Clinical Note Generation from Doctor-Patient Encounters (2023.eacl-main)

Copied to clipboard

Challenge: Medical doctors spend 52 to 102 minutes per day writing clinical notes from patient encounters.
Approach: They propose to use a new dataset to generate automated and manual clinical notes from doctor-patient conversations in a clinical setting.
Outcome: The proposed model could reduce the time spent writing clinical notes from doctor-patient conversations in a clinical setting.
To Err Is Human, How about Medical Large Language Models? Comparing Pre-trained Language Models for Medical Assessment Errors and Reliability (2024.lrec-main)

Copied to clipboard

Challenge: a 1999 report found that at least forty thousand deaths are a result of preventable medical errors.
Approach: They test pre-trained language models to characterize their error generation and reliability in medical assessment ability.
Outcome: The results show that pre-trained models can generate errors and perform better than human models.
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes (2025.findings-acl)

Copied to clipboard

Challenge: Several studies have shown that large language models can answer medical questions correctly, outperforming the average human score in some medical exams.
Approach: They introduce MEDEC, the first publicly available benchmark for medical error detection and correction in clinical notes.
Outcome: The proposed model outperforms medical doctors in errors detection and correction tasks.
Alignment Annotation for Clinic Visit Dialogue to Clinical Note Sentence Language Generation (2020.lrec-1)

Copied to clipboard

Challenge: Despite advances in natural language processing, converting a clinic visit conversation into a clinical note is a largely unexplored area of research.
Approach: They propose an annotation methodology that is content- and technique- agnostic while associating note sentences to sets of dialogue sentences.
Outcome: The proposed method is content- and technique-agnostic while associating note sentences to sets of dialogue sentences.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations