Challenge: HEAL is the first continuously trained LLaMA2-based LLM for medical conversations . despite the success of LLMs in general capabilities, they often fall short in niche domains like healthcare .
Approach: They propose a 13B LLaMA2-based LLM that is purpose-built for medical conversations and measured on automated scribing.
Outcome: The HEAL LLM outperforms GPT-4 and PMC-LLaMA in PubMedQA with 78.4% accuracy and parity with GPT-LLAMA in generating medical notes.

Similar Papers

Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation (2025.findings-acl)

Copied to clipboard

Challenge: Proprietary Large Language Models (LLMs) have demonstrated promising capabilities in clinical text summarization tasks.
Approach: They propose a domain- and task-specific adaptation process for an open-source LLaMA-2 model . LLama-2 can generate high-quality clinical notes from outpatient patient-doctor dialogues .
Outcome: The proposed model can generate clinical notes comparable to those authored by physicians.
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes (2025.findings-acl)

Copied to clipboard

Challenge: Several studies have shown that large language models can answer medical questions correctly, outperforming the average human score in some medical exams.
Approach: They introduce MEDEC, the first publicly available benchmark for medical error detection and correction in clinical notes.
Outcome: The proposed model outperforms medical doctors in errors detection and correction tasks.
LlamaCare: An Instruction Fine-Tuned Large Language Model for Clinical NLP (2024.lrec-main)

Copied to clipboard

Challenge: Large language models have shown remarkable abilities in generating natural texts . applying LLMs to clinical domain still poses significant challenges .
Approach: They propose a method of instruction fine-tuning for adapting large language models to clinical domains . they generate instructions, inputs, and outputs covering a wide spectrum of clinical services .
Outcome: The proposed method outperforms baseline LLMs on clinical tasks . it requires domain adaptation, task-specific learning, and reliability .
Comparing Two Model Designs for Clinical Note Generation; Is an LLM a Useful Evaluator of Consistency? (2024.findings-naacl)

Copied to clipboard

Challenge: a clinical note is a document that documents a doctor's interaction with a patient . authors show that LLMs can be used to measure quality indicators .
Approach: They analyze two different approaches to generate different sections of a SOAP note . they use PEGASUS-X Transformer models to examine note consistency .
Outcome: The proposed approach leads to similar ROUGE values and no difference in Factuality metric . human reviewers perform the same tasks with roughly the same agreement as the LLMs .
Don’t be my Doctor! Recognizing Healthcare Advice in Large Language Models (2024.emnlp-industry)

Copied to clipboard

Challenge: Large language models (LLMs) are becoming increasingly popular in everyday use, especially in highly regulated domains such as healthcare, where misleading advice may influence users to commit malpractice.
Approach: They present a large-scale health-advice benchmark dataset that evaluates large language models' ability to recognize health-related advice in industrial settings.
Outcome: The proposed model can be misinterpreted as direct advice in highly regulated domains such as healthcare, but the results are not enough to protect them from misinterpreting them as medical advice.
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Large-scale language models (LLMs) like ChatGPT have demonstrated impressive abilities in generating responses based on human instructions. however, their use in the medical domain can be challenging due to their lack of specific, in-depth knowledge.
Approach: They propose a system that integrates authoritative medical textbooks into LLMs’ framework using plug-and-play modules.
Outcome: The proposed system outperforms the specialized Med-PaLM 2 model on three medical QA tasks by 11.6% to 16.6%.
Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale (2024.emnlp-main)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) lack visual knowledge in medical applications due to data privacy concerns and high annotation costs.
Approach: They refined medical image-text pairs from PubMed and employed MLLMs (GPT-4V) to denoise and reformat the data.
Outcome: The proposed model significantly improves the MMMU Health & Medicine track and shows that it can be used in multimodal scenarios.
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation (2025.findings-emnlp)

Copied to clipboard

Challenge: Current medical benchmarks have limitations in question design, data sources and evaluation methods.
Approach: They propose a new benchmark covering five core medical areas . it includes 2,996 questions created from real-world electronic health records .
Outcome: The proposed model covers five core medical areas and includes 2,996 questions created from real-world electronic health records and expert-designed clinical scenarios.
MedINST: Meta Dataset of Biomedical Instructions (2024.findings-emnlp)

Copied to clipboard

Challenge: Medical data and tasks require extensive preprocessing and standardization for effective use in training LLMs.
Approach: They propose to use MedINST as a meta-dataset to evaluate LLMs' generalization ability.
Outcome: The meta-dataset of biomedical instruction measures the generalization ability of LLMs across multiple open-domain tasks.
Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: LLMs encode extensive knowledge within their parameters, but the knowledge in LLM models can become outdated over time.
Approach: They propose two new LLMs that provide outdated medical advice . they compare the models with a set of QA pairs whose verdict changed through time .
Outcome: The proposed models exhibit memorization of outdated knowledge to some extent.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations