Papers with MIMIC-III
Next Visit Diagnosis Prediction via Medical Code-Centric Multimodal Contrastive EHR Modelling with Hierarchical Regularisation (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies have not addressed the heterogeneous and hierarchical properties inherent in EHR data. |
| Approach: | They propose a medical code-centric multimodal contrastive EHR learning framework with hierarchical regularisation that integrates multifaceted information encompassing medical codes, demographics, and clinical notes. |
| Outcome: | The proposed framework integrates multifaceted information encompassing medical codes, demographics, and clinical notes using a tailored network design and bimodal contrastive losses. |
Writing habits and telltale neighbors: analyzing clinical concept usage patterns with sublanguage embeddings (D19-62)
Copied to clipboard
| Challenge: | Existing biomedical concepts may have multiple, often non-compositional surface forms, making them difficult to analyze using lexical occurrence alone. |
| Approach: | They propose a method for characterizing usage patterns of clinical concepts among different document types by embedding concepts on clinical documents of different types and measuring their nearest neighborhood structures. |
| Outcome: | Experiments on the MIMIC-III corpus show that the proposed method captures clinically relevant differences in concept usage while correcting for noise in embedding learning. |
Making the Most Out of the Limited Context Length: Predictive Power Varies with Clinical Note Type and Note Section (2023.acl-srw)
Copied to clipboard
| Challenge: | Clinical notes have a long time span over multiple long documents. |
| Approach: | They propose a framework to analyze clinical notes with high predictive power . they propose to combine different types of notes to improve performance . |
| Outcome: | The proposed framework could be used to extract information from clinical notes . it shows that the sample size can be optimized for large contexts . |
Ontological attention ensembles for capturing semantic concepts in ICD code prediction from clinical text (D19-62)
Copied to clipboard
Matus Falis, Maciej Pajak, Aneta Lisowska, Patrick Schrempf, Lucas Deckers, Shadia Mikhael, Sotirios Tsaftaris, Alison O’Neil
| Challenge: | a semantically interpretable system for automated ICD coding of clinical text documents is presented . coding errors may result in unpaid claims and loss of revenue, authors argue . |
| Approach: | They propose a semantically interpretable system for automated ICD coding of clinical text documents. |
| Outcome: | The proposed system improves on the MIMIC-III dataset by 2.7% relative to the previous state of the art. |
Toward Expanding the Scope of Radiology Report Summarization to Multiple Anatomies and Modalities (2023.acl-short)
Copied to clipboard
| Challenge: | Existing studies are limited to a single modality and a chest X-ray, making it difficult to replicate results or compare approaches. |
| Approach: | They propose a dataset to generate an impression section of a radiology report . they propose to use three new modalities and seven new anatomies to evaluate their models . |
| Outcome: | The proposed model is based on the MIMIC-III and MIMIC CXR datasets and evaluates their clinical efficacy via RadGraph, a factual correctness metric. |
THCM-CAL: Temporal-Hierarchical Causal Modelling with Conformal Calibration for Clinical Risk Prediction (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to risk prediction from EHRs handle structured diagnostic codes and unstructured narrative notes separately. |
| Approach: | They propose a Temporal-Hierarchical Causal Model with Conformal Calibration . they construct a multimodal causal graph where nodes represent clinical entities from two modalities . |
| Outcome: | The proposed model infers three clinically grounded interactions from textual propositions and ICD codes mapped to textual descriptions. |
CoPHE: A Count-Preserving Hierarchical Evaluation Metric in Large-Scale Multi-Label Text Classification (2021.emnlp-main)
Copied to clipboard
| Challenge: | Large-Scale Multi-Label Text Classification (LMTC) tasks with hierarchical label spaces include automatic assignment of ICD-9 codes to discharge summaries. |
| Approach: | They propose a set of metrics for hierarchical evaluation using the depth of the ontology to evaluate the predictions of neural LMTC models. |
| Outcome: | The proposed metrics compare with previous evaluations on prior art models for ICD-9 coding in MIMIC-III and propose further avenues of research involving the proposed representation. |
Does BERT Pretrained on Clinical Notes Reveal Sensitive Data? (2021.naacl-main)
Copied to clipboard
| Challenge: | Pretraining large (masked) language models over EHR data has yielded consistent performance gains across tasks. |
| Approach: | They propose to use large Transformers to release pretraining models over EHRs . they propose to recover patient names and conditions associated with them . |
| Outcome: | The proposed models recover patient names and conditions associated with patients . the proposed models share the model parameters for use by other researchers . |
Code Synonyms Do Matter: Multiple Synonyms Matching Network for Automatic ICD Coding (2022.acl-short)
Copied to clipboard
| Challenge: | Existing methods for automatic ICD coding use label attention to match related text snippets. |
| Approach: | They propose to use code synonyms to leverage for better code representation learning. |
| Outcome: | The proposed method outperforms previous state-of-the-art methods on the MIMIC-III dataset. |
CLIP: A Dataset for Extracting Action Items for Physicians from Hospital Discharge Notes (2021.acl-long)
Copied to clipboard
James Mullenbach, Yada Pruksachatkun, Sean Adler, Jennifer Seale, Jordan Swartz, Greg McKelvey, Hui Dai, Yi Yang, David Sontag
| Challenge: | Continuity of care is crucial to ensuring positive health outcomes for patients discharged from an inpatient hospital setting. |
| Approach: | They propose to annotate clinical action items from a dataset of medical notes annotated by physicians and extract them as multi-aspect extractive summarization. |
| Outcome: | The proposed dataset is annotated by physicians and covers 718 documents representing 100K sentences. |
Modelling Temporal Document Sequences for Clinical ICD Coding (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing studies on the ICD coding task focus on extracting codes from the discharge summary, but there is potential to automate the task by identifying relevant information from clinical notes. |
| Approach: | They propose a hierarchical transformer architecture that uses text across the entire sequence of clinical notes in each hospital stay for ICD coding. |
| Outcome: | The proposed model exceeds the state-of-the-art when using only discharge summaries as input and achieves performance improvements when all clinical notes are used as input. |
PM2F2N: Patient Multi-view Multi-modal Feature Fusion Networks for Clinical Outcome Prediction (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods focused on time series data but ignored clinical notes . fusion of multi-modal features of patients from different views is not feasible due to the time series and clinical notes data being stored as time series. |
| Approach: | They propose to combine time series and clinical notes to fuse multi-modal features of patients from different perspectives using graph neural networks. |
| Outcome: | The proposed method is superior to existing models on MIMIC-III benchmark. |
PromptEHR: Conditional Electronic Healthcare Records Generation with Prompt Learning (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for generating longitudinal multimodal EHRs are limited due to privacy concerns. |
| Approach: | They propose to generate longitudinal multimodal EHRs by unconditional generation or longitudinal inference . existing methods generate single-modal E HRs by conditional generation or by longitudinal inferment . |
| Outcome: | The proposed method is more flexible and controllable than existing methods and is more cost-effective than existing ones. |
Multi-label Few/Zero-shot Learning with Knowledge Aggregated from Multiple Label Graphs (2020.emnlp-main)
Copied to clipboard
| Challenge: | Few/zero-shot learning is a big challenge of many classification tasks, where a classifier is required to recognise instances of classes that have very few or even no training samples. |
| Approach: | They propose a multi-graph aggregation model that fuses knowledge from multiple label graphs encoding different semantic label relationships to improve multi-label zero/few-shot document classification. |
| Outcome: | The proposed model improves on two large clinical datasets and the EU legislation dataset on few/zero-shot labels. |
Multi-stage Retrieve and Re-rank Model for Automatic Medical Coding Recommendation (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for ICD indexing have a heavy label distribution and a manual process . Xie and Xing (2017) propose a new approach to ICD re-ranking . |
| Approach: | They propose a "retrieve and re-rank" framework to allocate subsets of ICD codes to medical records . they leverage auxiliary knowledge of the electronic health records (EHR) and a discrete retrieval method . |
| Outcome: | The proposed method achieves state-of-the-art performance on the MIMIC-III benchmark. |
On the Impact of Random Seeds on the Fairness of Clinical Classifiers (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent work has shown that fine-tuning large networks is surprisingly sensitive to changes in random seed(s). |
| Approach: | They explore the implications of this phenomenon for model fairness across demographic groups in clinical prediction tasks over electronic health records (EHR) they find that jointly optimizing for high overall performance and low disparities does not yield statistically significant improvements. |
| Outcome: | The proposed model fairness is based on the MIMIC-III dataset, the standard dataset in clinical NLP research. |
Clinical Note Owns its Hierarchy: Multi-Level Hypergraph Neural Networks for Patient-Level Representation Learning (2023.acl-long)
Copied to clipboard
| Challenge: | Clinical notes of patient EHRs contain valuable information from healthcare professionals, but have been underutilized due to their difficult-to-understand contents and complex hierarchies. |
| Approach: | They propose to use clinical notes to learn more balanced knowledge from EHRs by assembling useful neutral words with rare keywords via note and taxonomy level hyperedges. |
| Outcome: | The proposed method can retain clinical semantic information by (1) frequent neutral words and (2) hierarchies with imbalanced distribution. |
A New Public Corpus for Clinical Section Identification: MedSecId (2022.coling-1)
Copied to clipboard
| Challenge: | a study aims to segment sections of clinical medical domain documentation . section identification is a process by which sections are demarcated and labeled . |
| Approach: | They use a set of 2,002 fully annotated medical notes from the MIMIC-III to segment sections in clinical medical domain documentation. |
| Outcome: | The proposed model shows that medical concepts are related across sections using principal component analysis. |
Data Drift in Clinical Outcome Prediction from Admission Notes (2024.lrec-main)
Copied to clipboard
Paul Grundmann, Jens-Michalis Papaioannou, Tom Oberhauser, Thomas Steffek, Amy Siu, Wolfgang Nejdl, Alexander Loeser
| Challenge: | a pivotal dataset for clinical NLP research was released in 2016 . public access to such datasets is limited due to privacy and ethical concerns . |
| Approach: | They propose a novel clinical outcome prediction dataset based on MIMIC-IV . they provide initial insights into the performance of models trained on MIDIC-III . |
| Outcome: | The proposed dataset aims to probe the robustness and generalization of clinical outcome prediction models . the study focuses on challenges tied to evolving documentation standards and changing codes in the ICD taxonomy . |
A Cross-document Coreference Dataset for Longitudinal Tracking across Radiology Reports (2022.lrec-1)
Copied to clipboard
| Challenge: | Oftentimes, these findings and devices are referred to multiple times in a single report and are also referred across different reports of a patient. |
| Approach: | They propose a new cross-document coreference resolution (CDCR) dataset for identifying co-referring radiological findings and medical devices across a patient's radiology reports. |
| Outcome: | The proposed dataset contains 5872 mentions (findings and devices) spanning 638 MIMIC-III radiology reports across 60 patients, covering multiple imaging modalities and anatomies. |
Analyzing Code Embeddings for Coding Clinical Narratives (2021.findings-acl)
Copied to clipboard
| Challenge: | Recent work on automated ICD coding learn mappings between low-dimensional representations of clinical text reports and codes. |
| Approach: | They propose novel neural networks for encoding medical codes based on textual, structural and statistical characteristics using a single deep learning baseline model. |
| Outcome: | The proposed methods improve the accuracy of medical codes based on their textual, structural and statistical characteristics. |
THREAD: Thinking Deeper with Recursive Spawning (2025.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown impressive capabilities across diverse settings, but their performance degrades as context length and complexity increases. |
| Approach: | They propose to frame model generation as a thread of execution that, based on the context, can run to completion or dynamically spawn new threads. |
| Outcome: | The proposed model outperforms existing frameworks by 10% to 50% on diverse benchmarks. |
Effective Convolutional Attention Network for Multi-label Clinical Document Classification (2021.emnlp-main)
Copied to clipboard
| Challenge: | a large number of medical encounters need to be coded everyday due to long document sets and large label set. |
| Approach: | They propose a convolutional attention network for multi-label document classification problem . they use convolution-based encoders and convolution networks to aggregate information across documents . |
| Outcome: | The proposed model outperforms prior best model and multilingual Transformer model on a widely used dataset in the medical domain. |
CHiLL: Zero-shot Custom Interpretable Feature Extraction from Clinical Notes with Large Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study suggests that linear models with interpretable features are more reliable than opaque models. |
| Approach: | They propose an approach for natural-language specification of features for linear models . they prompt LLMs with expert-crafted queries to generate interpretable features from health records . |
| Outcome: | The proposed approach can be used to craft features clinically meaningful for downstream tasks . it is based on a risk prediction task and standard predictive tasks based upon this data . |
Text-Attributed Knowledge Graph Enrichment with Large Language Models for Medical Concept Representation (2026.acl-long)
Copied to clipboard
| Challenge: | eHRs encode a patient's medical history as a high-dimensional and sparse sequence of diagnosis, medication, and procedure concepts . robust concept representation learning is hindered by key challenges, authors say . clinically important cross-type dependencies are often missing or incomplete in existing ontology resources . |
| Approach: | They propose a graph learning framework that integrates semantics with medical concepts to improve prediction performance. |
| Outcome: | The proposed framework improves prediction performance and integrates semantics with graph structure. |
Learning What to Ignore: Mitigating Negative Transfer in Medical Knowledge Fusion via Clinical Task-Adaptive Selection (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to longitudinal EHR modeling struggle to balance structural authority of static ontologies with reasoning flexibility of large language models. |
| Approach: | They propose a framework that integrates external medical knowledge into longitudinal EHR modeling to mitigate clinical data sparsity. |
| Outcome: | The proposed framework outperforms state-of-the-art models on four clinical tasks. |
MedCPI: A Construct–Personalize–Integrate Framework for KG-enhanced Clinical Prediction (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing KG-enhanced approaches to clinical prediction are limited . existing approaches to personalize and integrate knowledge are weakly controlled . |
| Approach: | They propose a framework to integrate medical knowledge graphs into EHRs to support KG-enhanced clinical prediction. |
| Outcome: | The proposed framework improves on MIMIC-III and MIMIC IV tasks. |
No Black Boxes: Interpretable and Interactable Predictive Healthcare with Knowledge-Enhanced Agentic Causal Discovery (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Deep learning models lacking interpretability and interactivity, authors say . lack of interactive mechanisms prevents clinicians from incorporating their own knowledge into decision-making process. |
| Approach: | a new deep learning model is proposed to improve interpretability and interactivity . authors propose a knowledge-enhanced agent-driven causal discovery framework . |
| Outcome: | a new model improves interpretability and interactivity on EHR data . the proposed model improve interpretability through explicit reasoning and causal analysis . |
Learning Dynamic Representations and Policies from Multimodal Clinical Time-Series with Informative Missingness (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to accommodate missingness in clinical time series, but how to extract and use information carried by the observation process itself remains underexplored. |
| Approach: | They propose a patient representation learning framework that leverages informative missingness to learn multimodal clinical time series from structured and textual data. |
| Outcome: | The proposed framework improves offline treatment policy learning and adverse outcome prediction on ICU sepsis cohorts from MIMIC-III, MIMIC IV, and eICU. |
RePrompT: Recurrent Prompt Tuning for Integrating Structured EHR Encoders with Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown promising results for mining EHRs . translating time-stamped sequences into plain text can obscure both temporal structure and code identities, weakening the ability to capture code co-occurrence and longitudinal regularities. |
| Approach: | They propose a time-aware LLM framework that integrates structured EHR encoders through prompt tuning without modifying underlying architectures. |
| Outcome: | Experiments on MIMIC-III and MIMIC IV show that RePrompT outperforms both EHR-based and LLM-based baselines across multiple clinical prediction tasks. |