Papers by Wen-wai Yim
An Investigation of Evaluation Methods in Automatic Medical Note Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent studies show that doctors can save significant amounts of time when using automatic note generation. |
| Approach: | They propose task-specific metrics for automatic note generation from medical conversation summarization and generation, including knowledge-graph embedding-based metrics, customized model-based measures with domain-specific weights, and ensemble metrics. |
| Outcome: | The proposed evaluation metrics are compared to existing models and can have different behaviors on different types of clinical notes datasets. |
Automatic rubric-based content grading for clinical notes (D19-62)
Copied to clipboard
| Challenge: | Clinical notes are important documentation critical to medical care, as well as billing and legal needs. |
| Approach: | They propose to evaluate clinical note creation using rubric-based content grading . they build a feature-based system and a neural network-based baseline system . |
| Outcome: | The proposed system can be used to evaluate clinical notes for a rubric-based content grading system . the proposed system has content point accuracy and kappa values at 0.86 and 0.71 on the test set . |
An Empirical Study of Clinical Note Generation from Doctor-Patient Encounters (2023.eacl-main)
Copied to clipboard
| Challenge: | Medical doctors spend 52 to 102 minutes per day writing clinical notes from patient encounters. |
| Approach: | They propose to use a new dataset to generate automated and manual clinical notes from doctor-patient conversations in a clinical setting. |
| Outcome: | The proposed model could reduce the time spent writing clinical notes from doctor-patient conversations in a clinical setting. |
To Err Is Human, How about Medical Large Language Models? Comparing Pre-trained Language Models for Medical Assessment Errors and Reliability (2024.lrec-main)
Copied to clipboard
| Challenge: | a 1999 report found that at least forty thousand deaths are a result of preventable medical errors. |
| Approach: | They test pre-trained language models to characterize their error generation and reliability in medical assessment ability. |
| Outcome: | The results show that pre-trained models can generate errors and perform better than human models. |
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes (2025.findings-acl)
Copied to clipboard
| Challenge: | Several studies have shown that large language models can answer medical questions correctly, outperforming the average human score in some medical exams. |
| Approach: | They introduce MEDEC, the first publicly available benchmark for medical error detection and correction in clinical notes. |
| Outcome: | The proposed model outperforms medical doctors in errors detection and correction tasks. |
Alignment Annotation for Clinic Visit Dialogue to Clinical Note Sentence Language Generation (2020.lrec-1)
Copied to clipboard
| Challenge: | Despite advances in natural language processing, converting a clinic visit conversation into a clinical note is a largely unexplored area of research. |
| Approach: | They propose an annotation methodology that is content- and technique- agnostic while associating note sentences to sets of dialogue sentences. |
| Outcome: | The proposed method is content- and technique-agnostic while associating note sentences to sets of dialogue sentences. |