Challenge: Medical coding is time-consuming and error-prone due to large label space, lengthy text inputs, and the absence of supporting evidence annotations.
Approach: They propose a Generative AI framework for automatic medical coding that leverages extraction, retrieval, and re-ranking techniques as core components.
Outcome: The proposed framework outperforms existing methods on the International Classification of Diseases (ICD) code prediction scale.

Similar Papers

Code Like Humans: A Multi-Agent Solution for Medical Coding (2025.findings-emnlp)

Copied to clipboard

Challenge: In medical coding, experts map unstructured clinical notes to alphanumeric codes for diagnoses and procedures.
Approach: They introduce ‘Code Like Humans’: a new agentic framework for medical coding with large language models that implements official coding guidelines for human experts.
Outcome: The proposed framework implements official coding guidelines for human experts and can support the full ICD-10 coding system (+70K labels).
Aligning AI Research with the Needs of Clinical Coding Workflows: Eight Recommendations Based on US Data Analysis and Critical Review (2025.acl-long)

Copied to clipboard

Challenge: Clinical coding is labour-intensive and error-prone, which has motivated research towards full automation of the process.
Approach: They propose to use AI to improve evaluation methods and propose new methods to assist clinical coders in their workflows.
Outcome: The proposed methods can be improved and improved on existing methods and the existing ones to assist coders in their workflows.
A Neural Architecture for Automated ICD Coding (P18-1)

Copied to clipboard

Challenge: Medical coding is time-consuming, expensive, and error prone.
Approach: They propose to use diagnosis descriptions (DDs) of a patient as inputs to select the most relevant ICD codes.
Outcome: The proposed algorithms perform on a clinical dataset with 59K patient visits.
Scalable Wide and Deep Learning for Computer Assisted Coding (N18-3)

Copied to clipboard

Challenge: In recent years the use of electronic medical records has accelerated resulting in large volumes of medical data when a patient visits a healthcare facility.
Approach: They propose to use convolutional neural networks and logistic regression to build a machine learning based system for predicting ICD-10 codes from electronic medical records.
Outcome: The proposed system can predict ICD-10 codes from electronic medical records using convolutional neural networks and logistic regression models.
JointCoder: Exploring Automated ICD Coding on Real-World Chinese EHRs with a Multi-Agent Framework (2026.acl-demo)

Copied to clipboard

Challenge: Existing automated ICD coding systems face several fundamental challenges due to the limited availability of publicly available Chinese ICD datasets.
Approach: They propose to use a Chinese ICD coding dataset and a multi-agent framework to reformulate ICD as a joint disease-procedure coding task.
Outcome: The proposed system outperforms state-of-the-art methods on real-world Chinese ICD coding datasets and 1.7B-parameter models.
A Two-Stage Decoder for Efficient ICD Coding (2023.findings-acl)

Copied to clipboard

Challenge: Recent automated ICD coding efforts improve performance by encoding medical notes and codes with additional data and knowledge bases.
Approach: They propose a two-stage decoding mechanism to predict ICD codes using hierarchical properties of the codes to split the prediction into two steps: at first, predict the parent code and then predict the child code based on the previous prediction.
Outcome: Experiments on the public MIMIC-III data show that the proposed model performs well in single-model settings without external data or knowledge.
Less is More: Explainable and Efficient ICD Code Prediction with Clinical Entities (2025.acl-long)

Copied to clipboard

Challenge: Clinical coding is labor-intensive and prone to delays, leading to global backlogs.
Approach: They propose an approach that combines Named Entity Recognition (NER) and Assertion Classification (AC) to filter for clinically important content before supervised code prediction.
Outcome: The proposed approach reduces training time by over half on a standard evaluation dataset compared to current methods . it uses Named Entity Recognition (NER) and Assertion Classification (AC) to filter for clinically important content before supervised code prediction.
Auxiliary Knowledge-Induced Learning for Automatic Multi-Label Medical Document Classification (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for ICD indexing use machine learning to assign subset of codes to medical records . experimental results show proposed method achieves state-of-the-art performance on a number of measures.
Approach: They propose a method that uses a deep dilated residual convolution encoder to learn document representations across different lengths of the texts.
Outcome: The proposed method achieves state-of-the-art performance on a number of measures.
Analyzing Code Embeddings for Coding Clinical Narratives (2021.findings-acl)

Copied to clipboard

Challenge: Recent work on automated ICD coding learn mappings between low-dimensional representations of clinical text reports and codes.
Approach: They propose novel neural networks for encoding medical codes based on textual, structural and statistical characteristics using a single deep learning baseline model.
Outcome: The proposed methods improve the accuracy of medical codes based on their textual, structural and statistical characteristics.
Accurate and Well-Calibrated ICD Code Assignment Through Attention Over Diverse Label Embeddings (2024.eacl-long)

Copied to clipboard

Challenge: Existing approaches to assigning ICD codes to clinical text are time-consuming, labor intensive, and error-prone.
Approach: They propose to adapt a Transformer-based model to a longformer model and use it to encode clinical narratives.
Outcome: The proposed approach outperforms current state-of-the-art models in ICD coding with the label embeddings contributing to the good performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations