Papers by Ben Hachey

6 papers
Aligning AI Research with the Needs of Clinical Coding Workflows: Eight Recommendations Based on US Data Analysis and Critical Review (2025.acl-long)

Copied to clipboard

Challenge: Clinical coding is labour-intensive and error-prone, which has motivated research towards full automation of the process.
Approach: They propose to use AI to improve evaluation methods and propose new methods to assist clinical coders in their workflows.
Outcome: The proposed methods can be improved and improved on existing methods and the existing ones to assist coders in their workflows.
NNE: A Dataset for Nested Named Entity Recognition in English Newswire (P19-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is widely used in downstream tasks but most tools focus on flat mention structure over coarse schemas.
Approach: They describe a fine-grained, nested named entity dataset over the Wall Street Journal portion of the Penn Treebank.
Outcome: The proposed dataset comprises 279,795 mentions of 114 entity types with up to 6 layers of nesting.
Using Similarity Measures to Select Pretraining Data for NER (N19-1)

Copied to clipboard

Challenge: Existing studies on how to select appropriate data to pretrain word vectors or LMs are lacking.
Approach: They propose to quantify aspects of similarity between pretraining and target data.
Outcome: The proposed measures are good predictors of the usefulness of pretrained models for Named Entity Recognition over 30 data pairs.
An Effective Transition-based Model for Discontinuous NER (2020.acl-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) data sets often contain mentions consisting of discontinuous spans.
Approach: They propose a transition-based model with generic neural encoding for discontinuous NER that can recognize discontinuous mentions without sacrificing the accuracy on continuous mentions.
Outcome: The proposed model can recognize discontinuous mentions without sacrificing accuracy on continuous mentions.
Cost-effective Selection of Pretraining Data: A Case Study of Pretraining BERT on Social Media (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that domain-specific BERT models can be improved when in-domain data is used for pretraining.
Approach: They propose to use Twitter and forum text as pretraining sources for two BERT models and use similarity measures to nominate in-domain data for pretraining.
Outcome: The proposed method can be used to improve performance on downstream tasks by using in-domain data.
Less is More: Explainable and Efficient ICD Code Prediction with Clinical Entities (2025.acl-long)

Copied to clipboard

Challenge: Clinical coding is labor-intensive and prone to delays, leading to global backlogs.
Approach: They propose an approach that combines Named Entity Recognition (NER) and Assertion Classification (AC) to filter for clinically important content before supervised code prediction.
Outcome: The proposed approach reduces training time by over half on a standard evaluation dataset compared to current methods . it uses Named Entity Recognition (NER) and Assertion Classification (AC) to filter for clinically important content before supervised code prediction.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations