Challenge: In sensitive domains, the sharing of corpora is restricted due to confidentiality, copyrights or trade secrets.
Approach: They use auto-regressive neural models to generate a clinical case corpus annotated with clinical entities and evaluate it for a named entity recognition task.
Outcome: The proposed model can produce clinical case corpus annotated with clinical entities while maintaining confidentiality.

Similar Papers

Data-Constrained Synthesis of Training Data for De-Identification (2025.acl-long)

Copied to clipboard

Challenge: sensitive domains lack widely available datasets due to privacy risks . recent studies have focused on evaluating the privacy of the synthetic text .
Approach: They domain-adapt LLMs to clinical domain and generate synthetic clinical texts . they then generate NER models that can be annotated with tags for PII .
Outcome: The proposed model performs better than the original model using smaller datasets.
Named Entities in Medical Case Reports: Corpus and Experiments (2020.lrec-1)

Copied to clipboard

Challenge: Only very few annotated corpora in the medical domain exist.
Approach: They propose to annotate medical entities in case reports from PubMed Central's open access library.
Outcome: The proposed corpus is the first of its kind to be made available to the scientific community in English.
RecordTwin: Towards Creating Safe Synthetic Clinical Corpora (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to generate high-quality synthetic corpus from clinical documents require learning from the original clinical documents.
Approach: They propose a method to generate synthetic corpus from clinical documents using a large language model.
Outcome: The proposed method generates synthetic documents from in-hospital clinical documents.
Entity Decomposition with Filtering: A Zero-Shot Clinical Named Entity Recognition Framework (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have demonstrated that large language models (LLMs) can perform in named entity recognition tasks.
Approach: They propose a framework for clinical named entity recognition that decomposes the entity recognition task into several retrievals of sub-types and then filters them.
Outcome: The proposed framework improves on the clinical named entity recognition task.
Annotation of a Large Clinical Entity Corpus (D18-1)

Copied to clipboard

Challenge: Past researches have shown the superiority of statistical/ML approaches over the rule based approaches.
Approach: They propose to annotate a clinical domain annotated corpus using a small data set or a narrower domain to take full advantage of machine learning.
Outcome: The proposed corpus contains 5,160 clinical documents from forty different clinical specialties.
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes (2024.findings-acl)

Copied to clipboard

Challenge: Clinical notes are an extensive repository of information specific to individual patients.
Approach: They create synthetic large-scale clinical notes using publicly available case reports extracted from biomedical literature and train a clinical large language model, Asclepius.
Outcome: The proposed model outperforms several other models and is supported by detailed evaluations conducted by GPT-4 and medical professionals.
ClinicalT5: A Generative Language Model for Clinical Text (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent generative language models like BART and T5 are gaining popularity with their competitive performance on text generation and tasks cast as generative problems.
Approach: They propose to build domain-specific PLMs through fine-tuning or pre-training from scratch over domain corpora.
Outcome: The proposed model outperforms existing models on domain-specific tasks and compares favorably with its close baselines.
The Medical Scribe: Corpus Development and Model Performance Analyses (2020.lrec-1)

Copied to clipboard

Challenge: Existing tools to assist in clinical note generation using audio of provider-patient encounters are lacking.
Approach: They develop an annotation scheme to extract relevant clinical concepts from audio of provider-patient encounters and train a state-of-the-art tagging model.
Outcome: The proposed model is more useful than the F-scores reflect and can be used in clinical notes.
Named Entity Recognition for Chinese biomedical patents (2020.coling-main)

Copied to clipboard

Challenge: Existing attempts to address NER for Chinese biomedical texts have been limited due to the amount of Chinese biomedicine discoveries being patented.
Approach: They train and evaluate Chinese biomedical patents NER models based on BERT . their model is optimized for Chinese bio-patent data and scored an F1 .
Outcome: The proposed model achieves an F1 score of 0.540.15 for Chinese biomedical patent data.
Towards a Versatile Medical-Annotation Guideline Feasible Without Heavy Medical Knowledge: Starting From Critical Lung Diseases (2020.lrec-1)

Copied to clipboard

Challenge: Current annotation policies for medical corpora are not standardized across clinical texts of different types.
Approach: They propose to annotate medical records of various types using a named entity recognition (NER) task.
Outcome: The proposed annotation scheme is applicable to large-scale clinical NLP projects.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations