Challenge: a semantically interpretable system for automated ICD coding of clinical text documents is presented . coding errors may result in unpaid claims and loss of revenue, authors argue .
Approach: They propose a semantically interpretable system for automated ICD coding of clinical text documents.
Outcome: The proposed system improves on the MIMIC-III dataset by 2.7% relative to the previous state of the art.

Similar Papers

Less is More: Explainable and Efficient ICD Code Prediction with Clinical Entities (2025.acl-long)

Copied to clipboard

Challenge: Clinical coding is labor-intensive and prone to delays, leading to global backlogs.
Approach: They propose an approach that combines Named Entity Recognition (NER) and Assertion Classification (AC) to filter for clinically important content before supervised code prediction.
Outcome: The proposed approach reduces training time by over half on a standard evaluation dataset compared to current methods . it uses Named Entity Recognition (NER) and Assertion Classification (AC) to filter for clinically important content before supervised code prediction.
Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Current efforts in interpretability of medical coding rely heavily on label attention mechanisms, which often leads to the highlighting of extraneous tokens irrelevant to the ICD code.
Approach: They propose to leverage dictionary learning to extract sparsely activated representations from dense language models embedded in superposition to facilitate accurate interpretability.
Outcome: The proposed model extracts sparsely activated representations from dense language models in superposition, even when the highlighted tokens are medically irrelevant.
Evaluation and LLM-Guided Learning of ICD Coding Rationales (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies on the explainability of ICD coding rely on attention-based rationales and qualitative assessments conducted by physicians.
Approach: They propose to evaluate the explainability of rationales in ICD coding using a multi-granular rationale-annotated dataset.
Outcome: The proposed model improves the explainability of rationales in ICD coding by using human-annotated rationale-announced rationale models.
Clinical-Coder: Assigning Interpretable ICD-10 Codes to Chinese Clinical Notes (2020.acl-demos)

Copied to clipboard

Challenge: Existing methods of automatic coding prediction have been successful, but the interpretability of predicted codes is a challenge.
Approach: They propose an online system that can predict ICD codes for Chinese clinical notes by using a Dilated Convolutional Attention network with N-gram Matching mechanism.
Outcome: The proposed system is able to provide supporting information in clinical decision making.
Explainable Prediction of Medical Codes from Clinical Text (N18-1)

Copied to clipboard

Challenge: Clinical notes are text documents that are created by clinicians for each patient encounter.
Approach: They propose a method that aggregates information across the document using a convolutional neural network and uses an attention mechanism to select the most relevant segments for each of the thousands of possible codes.
Outcome: The proposed method is accurate and better than the current state of the art.
Fusion: Towards Automated ICD Coding via Feature Compression (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to assign ICD codes from unstructured clinical notes are noisy and prone to errors.
Approach: They propose a feature compressed ICD coding model called Fusion to address this problem.
Outcome: The proposed model outperforms existing models on two widely used datasets.
A Neural Architecture for Automated ICD Coding (P18-1)

Copied to clipboard

Challenge: Medical coding is time-consuming, expensive, and error prone.
Approach: They propose to use diagnosis descriptions (DDs) of a patient as inputs to select the most relevant ICD codes.
Outcome: The proposed algorithms perform on a clinical dataset with 59K patient visits.
ICDAGENT: Empowering Agentic Large Language Models for Explainable Medical Coding (2026.acl-long)

Copied to clipboard

Challenge: Existing models lack convincing, human-understandable explanations, making them difficult for physicians to trust and use in practice.
Approach: They propose a framework that aims to automatically assign ICD codes to clinical notes while providing explicit justifications for each assignment.
Outcome: The proposed framework achieves effective ICD coding with accurate explanations using two collaborative LLM agents: a coding agent and a critical agent.
Code Synonyms Do Matter: Multiple Synonyms Matching Network for Automatic ICD Coding (2022.acl-short)

Copied to clipboard

Challenge: Existing methods for automatic ICD coding use label attention to match related text snippets.
Approach: They propose to use code synonyms to leverage for better code representation learning.
Outcome: The proposed method outperforms previous state-of-the-art methods on the MIMIC-III dataset.
A General Knowledge Injection Framework for ICD Coding (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to improve ICD coding focus on a single type of knowledge and design specialized modules that are complex and incompatible with each other.
Approach: They propose a general knowledge injection framework that integrates three key types of knowledge without specialized design of additional modules.
Outcome: The proposed framework outperforms baseline models and is comparable to models relying on extra human annotations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations