Papers with EHR

49 papers
Next Visit Diagnosis Prediction via Medical Code-Centric Multimodal Contrastive EHR Modelling with Hierarchical Regularisation (2024.findings-eacl)

Copied to clipboard

Challenge: Existing studies have not addressed the heterogeneous and hierarchical properties inherent in EHR data.
Approach: They propose a medical code-centric multimodal contrastive EHR learning framework with hierarchical regularisation that integrates multifaceted information encompassing medical codes, demographics, and clinical notes.
Outcome: The proposed framework integrates multifaceted information encompassing medical codes, demographics, and clinical notes using a tailored network design and bimodal contrastive losses.
Applications of Natural Language Processing in Clinical Research and Practice (N19-5)

Copied to clipboard

Challenge: a tutorial on clinical NLP will introduce students and experts to the field . a focus will be on the use of clinical Nlp in clinical research and practice .
Approach: This tutorial introduces the clinical use of natural language processing (NLP) techniques . it will review techniques and tools developed for the clinical domain .
Outcome: This tutorial will introduce the clinical NLP methodologies and tools at two top universities . the goal of the tutorial is to encourage NLP researchers in the general domain to contribute .
FLIQA-AD: a Fusion Model with Large Language Model for Better Diagnose and MMSE Prediction of Alzheimer’s Disease (2025.naacl-short)

Copied to clipboard

Challenge: Existing classification and regression models that only extract finer-grained information from magnetic resonance imaging (MRI) may not be effective for Alzheimer's disease (AD).
Approach: They propose to use a 3D Adapter in a Vision Transformer to extract the patient's EHR information and questions related to the disease as text prompts.
Outcome: The proposed model can discriminate and predict the corresponding MMSE score based on the extracted brain structural information and textual content .
Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL (2026.eacl-long)

Copied to clipboard

Challenge: Despite recent advances, performance remains far from clinically reliable . specialized medical terminology and fine-grained temporal reasoning are key to executing clinical data analysis.
Approach: They propose a benchmark for clinical text-to-SQL that demands multi-table joins, clinically meaningful filters, and executable SQL.
Outcome: The proposed benchmark performs well on a set of 20 proprietary and open-source models . it scores 74.7% execution, while DeepSeek-R1 leads open-sourced at 69.2% .
Does BERT Pretrained on Clinical Notes Reveal Sensitive Data? (2021.naacl-main)

Copied to clipboard

Challenge: Pretraining large (masked) language models over EHR data has yielded consistent performance gains across tasks.
Approach: They propose to use large Transformers to release pretraining models over EHRs . they propose to recover patient names and conditions associated with them .
Outcome: The proposed models recover patient names and conditions associated with patients . the proposed models share the model parameters for use by other researchers .
ScAN: Suicide Attempt and Ideation Events Dataset (2022.naacl-main)

Copied to clipboard

Challenge: Suicidal behaviors, including suicide attempts (SA) and suicide ideations (SI), are leading risk factors for death by suicide.
Approach: They first built a Suicide Attempt and Ideation Events (ScAN) dataset, a subset of the publicly available MIMIC III dataset spanning over 12k+ EHR notes with 19k+ annotated SA and SI events information.
Outcome: The proposed model is based on the publicly available MIMIC III Suicide Attempt and Ideation Events Retriever (ScANER) dataset and achieves a macro-weighted F1 score of 0.83 for identifying suicidal behavioral evidences and a micro-weighting score of 0.8 and 0.60 for classification of SA and SI for the patient’s hospital-stay.
MedPlan: A Two-Stage RAG-Based System for Personalized Medical Plan Generation (2025.acl-industry)

Copied to clipboard

Challenge: Existing systems focus primarily on assessment rather than treatment planning.
Approach: They propose a framework that structures LLM reasoning to align with real-life workflows.
Outcome: The proposed framework outperforms baseline approaches in assessment accuracy and treatment plan quality.
DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries (2022.lrec-1)

Copied to clipboard

Challenge: a new question answering dataset is being developed for electronic health records . structured tables and unstructured notes can be duplicated, contradictory or provide additional context .
Approach: They develop a question-answer-matching dataset using structured tables and unstructured notes from an EHR.
Outcome: The proposed model is based on a model with a modality selection network . it uses the prediction of a RAT-SQL to choose between EHR tables and clinical notes .
Applications of BERT Models Towards Automation of Clinical Coding in Icelandic (2024.findings-naacl)

Copied to clipboard

Challenge: Traditionally, clinical coding is manual and laborintensive task prone to human error.
Approach: They analyze 25 years of electronic health records from the Landspitali University Hospital in Icelandic to explore the potential of using NLP for clinical coding.
Outcome: The best-performing model achieves competitive results in micro and macro F1 scores, with label attention contributing significantly to its success.
RuCCoD: Towards Automated ICD Coding in Russian (2025.emnlp-main)

Copied to clipboard

Challenge: a new dataset for clinical coding in Russian is available for download . human coders must navigate a wide array of medical terminology and time pressures .
Approach: They present a new dataset for ICD coding in Russian, a language with limited biomedical resources.
Outcome: The proposed model improves accuracy on an in-house EHR dataset from 2017 to 2021.
PM2F2N: Patient Multi-view Multi-modal Feature Fusion Networks for Clinical Outcome Prediction (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focused on time series data but ignored clinical notes . fusion of multi-modal features of patients from different views is not feasible due to the time series and clinical notes data being stored as time series.
Approach: They propose to combine time series and clinical notes to fuse multi-modal features of patients from different perspectives using graph neural networks.
Outcome: The proposed method is superior to existing models on MIMIC-III benchmark.
Hierarchical Pretraining on Multimodal Electronic Health Records (2023.emnlp-main)

Copied to clipboard

Challenge: Existing pretraining models on EHR data are too specific, limiting their transferability.
Approach: They propose a general, unified pretraining framework for hierarchically multimodal EHR data that can be used to train models on a large dataset before fine-tuning it on 'upstream' tasks.
Outcome: The proposed model performs on eight downstream tasks spanning three levels and compares with baselines on 18 different tasks.
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on static single-step calculations with explicit instructions.
Approach: They propose a benchmark for evaluating medical calculators in realistic scenarios . they use 118 scenario tasks across 4 clinical domains to evaluate medical calculator performance .
Outcome: The first benchmark for evaluating medical calculators in realistic scenarios is released . it features 118 scenario tasks across 4 clinical domains and is based on a model context protocol integration.
That’s the Wrong Lung! Evaluating and Improving the Interpretability of Unsupervised Multimodal Encoders for Medical Data (2022.emnlp-main)

Copied to clipboard

Challenge: Recent multimodal models induce soft local alignments between image regions and sentences.
Approach: They compare alignments from a state-of-the-art multimodal model for EHR with human annotations that link image regions to sentences.
Outcome: The proposed models induce soft local alignments between image regions and sentences . the text has an often weak or unintuitive influence on attention, the authors found .
ODD: A Benchmark Dataset for the Natural Language Processing Based Opioid Related Aberrant Behavior Detection (2024.naacl-long)

Copied to clipboard

Challenge: Opioid related aberrant behaviors (ORABs) present novel risk factors for opioid overdose.
Approach: They propose to use a biomedical natural language processing benchmark dataset to classify ORABs from patients’ EHR notes into nine categories: confirmed aberrant behavior, suggested aberrant behaviors, Opioids, indication, diagnosed opioid dependency, Benzodiazepines, medication changes, and Central Nervous System-related.
Outcome: The proposed dataset outperforms two state-of-the-art models in most categories and the gains are especially higher among uncommon classes.
Development of a Corpus Annotated with Medications and their Attributes in Psychiatric Health Records (2020.lrec-1)

Copied to clipboard

Challenge: Free text fields within electronic health records (EHRs) contain valuable clinical information which is often missed when conducting research using EHR databases.
Approach: They propose to extract medication annotations from mental health records by including contextual information around them.
Outcome: The aim of the study is to provide a more complete picture behind the mention of medications in the health records, by including additional contextual information around them.
Improving Neural Models for Radiology Report Retrieval with Lexicon-based Automated Annotation (2022.naacl-main)

Copied to clipboard

Challenge: Conventional exact or approximate termbased retrieval methods lack the ability of semantic understanding of the clinical as well as language context.
Approach: They combine clinical finding detection with supervised query match learning to train a model . findings are used as queries to train the Sentence-BERT model using triplet loss .
Outcome: The proposed method outperforms existing methods on multiple retrieval benchmarks.
Multi-stage Retrieve and Re-rank Model for Automatic Medical Coding Recommendation (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for ICD indexing have a heavy label distribution and a manual process . Xie and Xing (2017) propose a new approach to ICD re-ranking .
Approach: They propose a "retrieve and re-rank" framework to allocate subsets of ICD codes to medical records . they leverage auxiliary knowledge of the electronic health records (EHR) and a discrete retrieval method .
Outcome: The proposed method achieves state-of-the-art performance on the MIMIC-III benchmark.
On the Impact of Random Seeds on the Fairness of Clinical Classifiers (2021.naacl-main)

Copied to clipboard

Challenge: Recent work has shown that fine-tuning large networks is surprisingly sensitive to changes in random seed(s).
Approach: They explore the implications of this phenomenon for model fairness across demographic groups in clinical prediction tasks over electronic health records (EHR) they find that jointly optimizing for high overall performance and low disparities does not yield statistically significant improvements.
Outcome: The proposed model fairness is based on the MIMIC-III dataset, the standard dataset in clinical NLP research.
A Dual-Attention Network for Joint Named Entity Recognition and Sentence Classification of Adverse Drug Events (2020.findings-emnlp)

Copied to clipboard

Challenge: Adverse drug events (ADEs) are a leading cause of death in the United States and cost around $30 $130 billion every year.
Approach: They propose a multi-grained joint deep network to learn ADE entity recognition and ADE sentence classification tasks.
Outcome: The proposed model improves state-of-art F1 score on the MADE 1.0 benchmark of EHR notes.
When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications? (2024.findings-emnlp)

Copied to clipboard

Challenge: Numerical data is pivotal for medical questions and answers, but tabular data is not fully integrated into LLMs.
Approach: They examine the effectiveness of vector representations from last hidden states of LLMs for medical diagnostics and prognostics using electronic health record data.
Outcome: The proposed representations outperform those using raw numerical EHR data in medical diagnostics and prognostics.
How to leverage the multimodal EHR data for better medical prediction? (2021.emnlp-main)

Copied to clipboard

Challenge: Using deep learning to improve healthcare is challenging due to the complexity of EHR data.
Approach: They propose a method to integrate clinical notes from EHR and combine them with different data to improve prediction performance.
Outcome: The proposed model outperforms the state-of-the-art method without clinical notes on two prediction tasks.
Predicting in-hospital mortality by combining clinical notes with time-series data (2021.findings-acl)

Copied to clipboard

Challenge: In intensive care units, patient health is monitored through vital signals and clinical notes . previous work focused on predicting patient health using time-series data gathered from medical devices .
Approach: They propose a model that combines clinical notes and vital data to make accurate mortality predictions.
Outcome: The proposed model achieves an AUC score of 0.9, compared to the previous 0.87 . it can be used to make accurate in-hospital mortality predictions .
CTPD: Cross-Modal Temporal Pattern Discovery for Enhanced Multimodal Electronic Health Records Analysis (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for predicting clinical outcomes have focused on capturing temporal interactions within individual samples and fusing multimodal information, overlooking critical temporal patterns across different patients.
Approach: They propose a cross-modal temporal pattern discovery framework to extract temporal patterns from multimodal EHR data.
Outcome: The proposed framework extracts meaningful cross-modal temporal patterns from multimodal EHR data.
Leveraging Medical Literature for Section Prediction in Electronic Health Records (D19-1)

Copied to clipboard

Challenge: Prior approaches to section prediction have only used text data from EHRs and required significant manual annotation.
Approach: They propose to use sections from medical literature to train models to predict sections in EHRs.
Outcome: The proposed model uses sections from medical literature that contain similar content to those found in EHR sections.
HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that LLM-based EHR question answering is costly to deploy and does not leverage hierarchical structure of clinical data.
Approach: They propose a Lorentzian model that embeds codes, visits, and questions in hyperbolic space and answers queries via geometry-consistent cross-attention with type-specific pointer heads.
Outcome: The proposed model embeds codes, visits, and questions in hyperbolic space and answers queries via geometry-consistent cross-attention with type-specific pointer heads.
Generation of Patient After-Visit Summaries to Support Physicians (2022.coling-1)

Copied to clipboard

Challenge: After-visit summary is a summary note given to patients after their clinical visit.
Approach: They propose to automate the generation of after-visit summaries and introduce a feedback mechanism that alerts physicians when an automatic summary fails to capture important details of the clinical notes.
Outcome: The proposed system improves on a large clinical dataset that contains electronic health record (EHR) notes and their associated summaries.
A Semi-supervised Approach for De-identification of Swedish Clinical Text (2020.lrec-1)

Copied to clipboard

Challenge: An abundance of electronic health records (EHRs) is produced every day within healthcare.
Approach: They propose a semi-supervised method for automatically creating high-quality training data for de-identification using annotated data for training and annotations that are costly in time and human resources.
Outcome: The proposed method improves recall from 84.75% to 89.20% without sacrificing precision to the same extent, dropping from 95.73% to 94.20%.
CARER - ClinicAl Reasoning-Enhanced Representation for Temporal Health Risk Prediction (2024.emnlp-main)

Copied to clipboard

Challenge: Existing deep learning methods require large datasets to achieve high generalizability.
Approach: They propose a framework that enhances deep learning models with clinical rationales derived from medically proficient Large Language Models.
Outcome: The proposed framework outperforms state-of-the-art models on two tasks using two popular EHR datasets by up to 11.2%.
Adversarial Learning of Privacy-Preserving Text Representations for De-Identification of Medical Records (P19-1)

Copied to clipboard

Challenge: De-identification is the task of detecting protected health information (PHI) in medical text.
Approach: They propose to create shareable representations of medical text that contain no PHI and can be shared between organizations to create unified datasets for training de-identification models.
Outcome: The proposed representation allows training a simple LSTM-CRF model to an F1 score of 97.4%.
Hierarchical Annotation for Building A Suite of Clinical Natural Language Processing Tasks: Progress Note Understanding (2022.lrec-1)

Copied to clipboard

Challenge: Existing corpus and annotations focus on textual features and relation prediction, but there are no structured corpus models for clinical diagnostic thinking.
Approach: They propose a hierarchical annotation schema with three stages to address clinical diagnostic thinking.
Outcome: The proposed model is based on a large collection of publicly available daily progress notes.
Extracting Biomedical Entities from Noisy Audio Transcripts (2024.lrec-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is particularly affected by noise, often termed the ASR-NLP gap.
Approach: They propose a dataset to bridge the ASR-NLP gap in the biomedical domain by extracting adverse drug reactions and mentions of entities from the Brief Test of Adult Cognition by Telephone (BTACT) exam.
Outcome: The proposed method can clean 2,000 clean and noisy recordings and eliminate errors using zero-shot and few-shot methods.
Extracting Social Determinants of Health from Pediatric Patient Notes Using Large Language Models: Novel Corpus and Methods (2024.lrec-main)

Copied to clipboard

Challenge: Social determinants of health (SDoH) are often studied in the electronic health record (EHR) however, there are difficulties in documenting SDoH in a tabular format due to the lack of a comprehensive SDoh tool.
Approach: They propose to annotate social history sections from 1,260 clinical notes from pediatric patients within the University of Washington (UW) hospital system.
Outcome: The proposed corpus captures ten distinct health determinants including living and economic stability, prior trauma, education access, substance use history, and mental health with an overall annotator agreement of 81.9 F1.
DKEC: Domain Knowledge Enhanced Multi-Label Classification for Diagnosis Prediction (2024.emnlp-main)

Copied to clipboard

Challenge: Prior work focused on hierarchical label structures but neglected to incorporate external knowledge from medical guidelines.
Approach: They propose to incorporate external knowledge from medical guidelines into domain knowledge enhanced classification for diagnosis prediction.
Outcome: The proposed system outperforms state-of-the-art label-wise attention networks and transformer models on a real-world emergency medical services dataset and a public electronic health record dataset.
Dataset and Enhanced Model for Eligibility Criteria-to-SQL Semantic Parsing (2020.lrec-1)

Copied to clipboard

Challenge: Clinical trials require that patients meet eligibility criteria to ensure safety and effectiveness of studies.
Approach: They propose a dataset that includes the first-of-its-kind eligibility-criteria corpus and queries for criteria-to-sql . they propose 'neuro semantic parser' which can translate eligibility criteria to executable SQL queries .
Outcome: The proposed parser outperforms existing state-of-the-art general-purpose models while highlighting the challenges presented by the new dataset.
README: Bridging Medical Jargon and Lay Understanding for Patient Education through Data-Centric NLP (2024.findings-emnlp)

Copied to clipboard

Challenge: a new task is to generate lay definitions of medical terms in EHRs that are difficult to understand for patients.
Approach: They propose a task of automatically generating lay definitions to simplify medical terms into patient-friendly lay language.
Outcome: The proposed model can match or surpass state-of-the-art closed-source large language models like ChatGPT with high-quality data.
Text-Attributed Knowledge Graph Enrichment with Large Language Models for Medical Concept Representation (2026.acl-long)

Copied to clipboard

Challenge: eHRs encode a patient's medical history as a high-dimensional and sparse sequence of diagnosis, medication, and procedure concepts . robust concept representation learning is hindered by key challenges, authors say . clinically important cross-type dependencies are often missing or incomplete in existing ontology resources .
Approach: They propose a graph learning framework that integrates semantics with medical concepts to improve prediction performance.
Outcome: The proposed framework improves prediction performance and integrates semantics with graph structure.
MedJEx: A Medical Jargon Extraction Model with Wiki’s Hyperlink Span and Contextualized Masked Language Model Score (2022.emnlp-main)

Copied to clipboard

Challenge: Existing natural language processing (NLP) methods for identifying medical jargon terms are difficult for patients to understand.
Approach: They propose a natural language processing application for identifying medical jargon terms from electronic health record notes.
Outcome: The proposed model outperforms state-of-the-art models on an auxiliary Wikipedia hyperlink span dataset and on the annotated MedJ dataset.
EHR-SeqSQL : A Sequential Text-to-SQL Dataset For Interactively Exploring Electronic Health Records (2024.findings-acl)

Copied to clipboard

Challenge: EHR-SeqSQL is the first text-to-SQl dataset to include sequential and contextual questions.
Approach: They propose a sequential text-to-SQL dataset for electronic health records databases that addresses critical yet underexplored aspects in text- to-SqL parsing.
Outcome: The proposed dataset improves compositional generalization efficiency and improves interactivity and compositionality.
MHGRL: An Effective Representation Learning Model for Electronic Health Records (2024.lrec-main)

Copied to clipboard

Challenge: Effective EHR representations are key to achieving high performance in healthcare applications.
Approach: They propose a multimodal heterogeneous graph-enhanced representation learning to learn EHR representations using medical ontology and textual notes.
Outcome: The proposed model outperforms baseline models on two real clinical datasets in downstream tasks.
ReMedi: Reasoner for Medical Clinical Prediction (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to predicting future clinical outcomes from EHRs focus on enhancing medical knowledge through distillation or RAG while relying on the model’s internal ability to interpret contextual information.
Approach: They propose a framework for improving clinical outcome prediction from EHR using a sample regeneration mechanism that leverages ground-truth answers as hints to enhance reasoning.
Outcome: Experiments on multiple EHR prediction tasks show significant gains of up to 19.9% over state-of-the-art baselines in terms of F1 score, underscoring ReMedi’s effectiveness in real-world clinical prediction.
Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QA (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to improve the reliability of Large Language Models (LLMs) in clinical applications require factual knowledge from open-ended datasets and clinical case-based knowledge to provide context grounded in real-world patient experiences.
Approach: They propose a retrieval-augmented generation framework based on the electronic health record to offer contextual information from other patients’ discharge reports.
Outcome: The proposed framework outperforms a text-based ranker in a clinical QA dataset with 1,280 discharge-related questions .
Large Language Models with Temporal Reasoning for Longitudinal Clinical Summarization and Prediction (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models have shown potential in clinical text summarization, but their ability to handle long patient trajectories with multi-modal data spread across time remains underexplored.
Approach: They evaluate open-source large language models, their Retrieval Augmented Generation variants and chain-of-thought prompting on long-context clinical summarization and prediction.
Outcome: The proposed models can synthesize structured and unstructured EHR data while reasoning over temporal coherence.
MedCPI: A Construct–Personalize–Integrate Framework for KG-enhanced Clinical Prediction (2026.findings-acl)

Copied to clipboard

Challenge: Existing KG-enhanced approaches to clinical prediction are limited . existing approaches to personalize and integrate knowledge are weakly controlled .
Approach: They propose a framework to integrate medical knowledge graphs into EHRs to support KG-enhanced clinical prediction.
Outcome: The proposed framework improves on MIMIC-III and MIMIC IV tasks.
Follow-up Question Generation For Enhanced Patient-Provider Conversations (2025.acl-long)

Copied to clipboard

Challenge: Follow-up question generation is an essential feature of dialogue systems as it can reduce conversational ambiguity and enhance modeling complex interactions.
Approach: They propose a framework that generates personalized follow-up questions based on patient utterances and prior EHR data.
Outcome: The framework reduces follow-up communications by 34% and improves performance by 17% and 5% on real and synthetic data.
EHRAgent: Code Empowers Large Language Models for Few-shot Complex Tabular Reasoning on Electronic Health Records (2024.emnlp-main)

Copied to clipboard

Challenge: EHRAgent enables clinicians to interact with EHRs using natural language . reliance on rule-based conversion systems often necessitates additional training or effort from data engineers.
Approach: They propose a large language model agent that generates and executes code in natural language to facilitate clinicians in directly interacting with EHRs.
Outcome: The proposed agent outperforms the strongest baseline by up to 29.6% in success rate on three real-world EHR datasets.
No Black Boxes: Interpretable and Interactable Predictive Healthcare with Knowledge-Enhanced Agentic Causal Discovery (2025.findings-emnlp)

Copied to clipboard

Challenge: Deep learning models lacking interpretability and interactivity, authors say . lack of interactive mechanisms prevents clinicians from incorporating their own knowledge into decision-making process.
Approach: a new deep learning model is proposed to improve interpretability and interactivity . authors propose a knowledge-enhanced agent-driven causal discovery framework .
Outcome: a new model improves interpretability and interactivity on EHR data . the proposed model improve interpretability through explicit reasoning and causal analysis .
DiaLLMs: EHR-Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction (2025.findings-acl)

Copied to clipboard

Challenge: Existing medical LLMs focus primarily on diagnosis recommendation, limiting their clinical applicability.
Approach: They propose a medical LLM that integrates heterogeneous EHR data into clinically grounded dialogues.
Outcome: The proposed model outperforms baselines in clinical test recommendation and diagnosis prediction.
RePrompT: Recurrent Prompt Tuning for Integrating Structured EHR Encoders with Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown promising results for mining EHRs . translating time-stamped sequences into plain text can obscure both temporal structure and code identities, weakening the ability to capture code co-occurrence and longitudinal regularities.
Approach: They propose a time-aware LLM framework that integrates structured EHR encoders through prompt tuning without modifying underlying architectures.
Outcome: Experiments on MIMIC-III and MIMIC IV show that RePrompT outperforms both EHR-based and LLM-based baselines across multiple clinical prediction tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations