Papers with PubMed

46 papers
BiomedCurator: Data Curation for Biomedical Literature (2022.aacl-demo)

Copied to clipboard

Challenge: BiomedCurator uses state-of-the-art natural language processing techniques to extract structured data from scientific articles.
Approach: They propose a web application that extracts structured data from PubMed and ClinicalTrials.gov . the application uses a combination of natural language processing techniques and a pattern-based extraction approach .
Outcome: The proposed system extracts the structured data from PubMed and ClinicalTrials.gov datasets.
EVIDENCEMINER: Textual Evidence Discovery for Life Sciences (2020.acl-demos)

Copied to clipboard

Challenge: EVIDENCEMINER is a web-based system that allows users to query a natural language statement and retrieve textual evidence from a background corpora for life sciences.
Approach: They propose a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences.
Outcome: EVIDENCEMINER is a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences.
Trigger Word Detection and Thematic Role Identification via BERT and Multitask Learning (D19-57)

Copied to clipboard

Challenge: Using natural language processing to discover and mine drug-related knowledge from text has been a hot topic in recent years.
Approach: They propose to use a pre-trained biomedical language representation model to extract mutation-disease knowledge from PubMed.
Outcome: The proposed approaches achieve 0.60 (ranks 1) and 0.25 (rank 2) on task 1 and task 2 respectively in terms of F1 metric.
SciClaims: An End-to-End Generative System for Biomedical Claim Analysis (2025.emnlp-demos)

Copied to clipboard

Challenge: SciClaims is an interactive web-based system for scientific claim analysis in the biomedical domain.
Approach: They present SciClaims, an interactive web-based system for scientific claim analysis in the biomedical domain.
Outcome: The system extracts factual claims from scientific texts and retrieves evidence from PubMed . it also verifies the validity of each claim using large language models . the system is optimized to run efficiently on a single GPU and is publicly available .
A Multi-Task Learning Framework for Extracting Bacteria Biotope Information (D19-57)

Copied to clipboard

Challenge: Existing methods to extract information from unstructured text are slow or expensive to get.
Approach: They propose a multi-task transfer multi-learning method for Bacteria Biotope rel+ner task . they use BERT and pre-train it using mask language models and next sentence prediction .
Outcome: The proposed method achieves the best performance on all metrics including slot error rate, precision and recall in the Bacteria Biotope rel+ner subtask.
Content-Based Citation Recommendation (N18-1)

Copied to clipboard

Challenge: Existing citation recommendation systems rely on information of query documents such as author names and publication venue.
Approach: They propose a content-based method for recommending citations in academic paper drafts . they embed a given query document into a vector space and use its nearest neighbors as candidates .
Outcome: The proposed method outperforms published methods on PubMed and DBLP datasets without metadata.
MedTutor: A Retrieval-Augmented LLM System for Case-Based Medical Education (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing educational tools for medical residents are time-consuming and inconsistent.
Approach: They propose a system that generates educational content and multiple-choice questions from clinical case reports and a pipeline that takes clinical case report input and produces targeted educational materials.
Outcome: The system generates educational content and multiple-choice questions from clinical case reports and synergizes with local knowledge base to ensure it is foundationally sound and current.
WikiSum: Coherent Summarization Dataset for Efficient Human-Evaluation (2021.acl-short)

Copied to clipboard

Challenge: Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems .
Approach: They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language.
Outcome: The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature.
LoRaLay: A Multilingual and Multimodal Dataset for Long Range and Layout-Aware Summarization (2023.eacl-main)

Copied to clipboard

Challenge: Text Summarization is a popular task and a challenge for neural models.
Approach: They propose to exploit visual/layout information to capture long-range dependencies in summarization models by combining layout-aware and long-reaching models.
Outcome: The proposed datasets cover French, Spanish, Portuguese, and Korean languages.
Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable Text (2025.naacl-industry)

Copied to clipboard

Challenge: Proteins play critical roles in biological systems, yet 99.7% of 227 million known protein sequences remain uncharacterized due to the limitations of experimental methods.
Approach: They propose a multimodal large language model that interprets protein sequences and generates informative text to address open-ended questions about protein functions and attributes.
Outcome: The proposed model outperforms existing models in open-ended question-answering tasks.
Sequential Span Classification with Neural Semi-Markov CRFs for Biomedical Abstracts (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for dividing biomedical abstracts into rhetorical segments assign a rhetorical label to each sentence while considering context in the abstract.
Approach: They propose to use Neural Semi-Markov Conditional Random Fields to assign a rhetorical label to a span that consists of continuous sentences.
Outcome: The proposed method achieved the best micro sentence-F1 score and the best macro span-F1.
HiGen: Hierarchy-Aware Sequence Generation for Hierarchical Text Classification (2024.eacl-long)

Copied to clipboard

Challenge: Hierarchical text classification is a complex subtask under multi-label text classification . the relevance of document sections can vary based on the hierarchy level, necessitating a dynamic document representation.
Approach: They propose a text-generation-based framework that uses language models to encode dynamic text representations.
Outcome: The proposed framework surpasses existing methods while handling data and mitigating class imbalance.
OTExtSum: Extractive Text Summarisation with Optimal Transport (2022.findings-naacl)

Copied to clipboard

Challenge: Extractive text summarisation aims to select salient sentences from a document to form a short yet informative summary.
Approach: They propose to formulate extractive text summarisation as an Optimal Transport (OT) problem and use it to obtain an optimal summary that minimises the transportation cost to a given document.
Outcome: The proposed method outperforms state-of-the-art methods and learning-based methods on multiNews, PubMed, BillSum, and CNN/DM datasets.
PubMedCLIP: How Much Does CLIP Benefit Visual Question Answering in the Medical Domain? (2023.findings-eacl)

Copied to clipboard

Challenge: Medical visual question answering is a multimodal task that requires a system to understand both medical images and textual questions and infer associations between them.
Approach: They propose a fine-tuned version of CLIP for the medical domain based on PubMed articles.
Outcome: The proposed model improves accuracy up to 3% on two MedVQA benchmark datasets.
Discourse-Aware Unsupervised Summarization for Long Scientific Documents (2021.eacl-main)

Copied to clipboard

Challenge: Existing extractive models for short news summarization are weak, despite recent advances in abstractive summarizing.
Approach: They propose an unsupervised graph-based ranking model that uses a hierarchical graph representation to determine sentence importance.
Outcome: The proposed model outperforms strong unsupervised baselines by wide margins in automatic metrics and human evaluation.
HiStruct+: Improving Extractive Text Summarization with Hierarchical Structure Information (2022.findings-acl)

Copied to clipboard

Challenge: Existing models that treat texts as linear sequences do not include hierarchical structure information.
Approach: They propose to inject hierarchical structure information into an extractive summarization model by combining hierarchically structured text with a pre-trained Transformer language model.
Outcome: The proposed model outperforms a baseline model on PubMed and arXiv datasets and the hierarchical structure information is not injected.
Efficient Attentions for Long Document Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Existing models that use full attentions have quadratic computational and memory complexities, and are too costly for long documents.
Approach: They propose an efficient encoder-decoder attention with head-wise positional strides to effectively pinpoint salient information from the source.
Outcome: The proposed model can process ten times more tokens than current models that use full attentions.
LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization (2023.eacl-main)

Copied to clipboard

Challenge: Human evaluation is labor-intensive, expensive to scale, and difficult to design.
Approach: They propose a set of guidelines for human evaluation of faithfulness in long-form summaries that address the following challenges: (1) How can we achieve high inter-annotator agreement on faithfulness scores? (2) How can our annotator minimize workload while maintaining accurate faithfulness?
Outcome: The proposed framework reduces inter-annotator variance in faithfulness scores while minimizing annotator workload while maintaining accuracy.
Comparing Knowledge Sources for Open-Domain Scientific Claim Verification (2024.eacl-long)

Copied to clipboard

Challenge: Existing systems for fact-checking scientific claims assume that the documents containing the evidence are already provided and annotated or contained in a limited corpus.
Approach: They perform an array of experiments to test the performance of open-domain claim verification systems on four datasets of biomedical and health claims in different settings.
Outcome: The proposed system performs better with biomedical and health claims, while Wikipedia is more suited for everyday health concerns.
Article Classification with Graph Neural Networks and Multigraphs (2024.lrec-main)

Copied to clipboard

Challenge: Existing and newly published articles require complex and complex pipelines to classify them into context-specific label taxonomies.
Approach: They propose to enrich Graph Neural Network pipelines with multi-graph representations that encode multiple signals of article relatedness as distinct edge types.
Outcome: The proposed methods improve the performance of a variety of GNN models compared to default graphs.
MolXPT: Wrapping Molecules with Text for Generative Pre-training (2023.acl-short)

Copied to clipboard

Challenge: Experimental results show that Generative pre-trained Transformers (GPT) have great success in natural language processing.
Approach: They propose a unified language model of text and molecules pre-trained on SMILES wrapped by text.
Outcome: The proposed model outperforms strong baselines of molecular property prediction on MoleculeNet and performs comparably to the best model in text-molecule translation while using less than half of its parameters.
“According to . . . ”: Prompting Language Models Improves Quoting from Pre-Training Data (2024.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) may hallucinate and generate false information despite pre-training on factual data.
Approach: They propose a new evaluation metric that measures the extent to which model-produced answers are directly found in underlying text corpora.
Outcome: The proposed evaluation metric measures the extent to which model-produced answers are directly found in underlying text corpora.
StrucSum: Graph-Structured Reasoning for Long Document Extractive Summarization with LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) have shown strong performance in zero-shot summarization, but struggle to model document structure and identify salient information in long texts.
Approach: They propose a training-free prompting framework that injects structural signals into prompts via sentence-level graph structures.
Outcome: The proposed framework improves summary quality and factual consistency over baselines and vanilla prompting.
Sparse Parallel Training of Hierarchical Dirichlet Process Topic Models (2020.emnlp-main)

Copied to clipboard

Challenge: To scale non-parametric extensions of probabilistic topic models, practitioners rely increasingly on parallel and distributed systems.
Approach: They propose a data-parallel sampler that utilizes all available sources of sparsity found in natural language to control memory requirements and computational complexity.
Outcome: The proposed sampler is able to train a hierarchical Dirichlet process topic model on a well-known corpus (PubMed) with 8m documents and 768m tokens, using a single multi-core machine in under four days.
Judge and Improve: Towards a Better Reasoning of Knowledge Graphs with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to integrating graph and language models face two key limitations: achieving robust semantic alignment and ensuring interpretability in outputs.
Approach: They propose a framework to integrate graph and language modalities while enhancing transparency.
Outcome: Extensive experiments on three benchmark datasets show that the proposed framework surpasses existing methods in efficiency and generates outputs that are significantly more interpretable.
Improving Health Question Answering with Reliable and Time-Aware Evidence Retrieval (2024.findings-naacl)

Copied to clipboard

Challenge: Existing question answering systems rely on pre-selected and annotated evidence documents, thus making them inadequate for addressing novel questions.
Approach: They propose to use the common retrieve-then-read QA pipeline and PubMed as a trustworthy collection of medical research documents to answer health questions from three diverse datasets.
Outcome: The proposed approach improves the macro F1 score by 10% by utilizing the common retrieve-then-read QA pipeline and PubMed as a trustworthy collection of medical research documents.
Extracting Fine-Grained Knowledge Graphs of Scientific Claims: Dataset and Transformer-Based Results (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches focus on high-level description of how research is carried out . instead, we focus on the subtleties of how experimental associations are presented .
Approach: They propose a transformer-based approach to relational scientific information extraction that captures associations over experimental variables and their qualifications, subtypes, and evidence.
Outcome: The proposed schema captures causal, comparative, predictive, statistical, and proportional associations over experimental variables along with qualifications, subtypes, and evidence.
BioReader: a Retrieval-Enhanced Text-to-Text Transformer for Biomedical Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Recent research has equipped language models with the ability to attend over relevant and factual information from non-parametric external sources, drawing a complementary path to architectural scaling.
Approach: They propose a retrieval-enhanced text-to-text model that augments the input prompt by fetching and assembling relevant scientific literature chunks from a neural database centered on PubMed.
Outcome: The proposed model outperforms state-of-the-art models on a broad array of downstream tasks while using up to 3x fewer parameters.
Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale (2024.emnlp-main)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) lack visual knowledge in medical applications due to data privacy concerns and high annotation costs.
Approach: They refined medical image-text pairs from PubMed and employed MLLMs (GPT-4V) to denoise and reformat the data.
Outcome: The proposed model significantly improves the MMMU Health & Medicine track and shows that it can be used in multimodal scenarios.
Factorizing Content and Budget Decisions in Abstractive Summarization of Long Documents (2022.emnlp-main)

Copied to clipboard

Challenge: Using a factorization approach, summarization decisions are conflated into a single feedforward step without taking into account contextual factors.
Approach: They propose to factorize summarization into two steps following a budget and content guidance.
Outcome: The proposed method outperforms PEGASUS in domain adaptation and generates significantly higher ROUGE scores on multiple benchmarks for long document summarization.
MemSum: Extractive Summarization of Long Documents Using Multi-Step Episodic Markov Decision Processes (2022.acl-long)

Copied to clipboard

Challenge: MemSum is a reinforcement-learning-based extractive summarizer that considers the text content of the sentence, the global context of the rest of the document, and the extraction history of the sentences that have already been extracted.
Approach: They propose a reinforcement-learning-based extractive summarizer that iteratively selects sentences from a broad set of information that would intuitively be used by humans.
Outcome: The proposed extractive summarizer is enriched with information on the extraction history and local, global, and historical information.
MEDLINE as a Parallel Corpus: a Survey to Gain Insight on French-, Spanish- and Portuguese-speaking Authors’ Abstract Writing Practice (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpora are used to train and evaluate machine translation systems, but little information is available about the methods used for producing the corpus, including translation direction.
Approach: They used PubMed and publisher websites to obtain contact information for MEDLINE authors and asked about their abstract writing practices.
Outcome: The authors of MEDLINE articles included in the English/Spanish, English/FR, and English/Portuguese (EN/PT) WMT 2019 test sets reported a response rate of over 20% .
Document Set Expansion with Positive-Unlabeled Learning Using Intractable Density Estimation (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to Document Set Expansion (DSE) rely on the unrealistic assumption of knowing the class prior for positive samples in the collection.
Approach: They propose a novel method that utilizes intractable density estimation models to learn the class prior for positive samples in the collection.
Outcome: The proposed method is based on a set of examples from PubMed and Covid datasets in a transductive setting.
Detecting Causal Language Use in Science Findings (D19-1)

Copied to clipboard

Challenge: Prior studies on identifying inappropriate use of causal language relied on manual content analysis, which is not scalable for examining a large volume of science publications.
Approach: They developed a prediction model that classifies conclusion sentences into “no relationship”, “correlational”, “conditional causal” and “direct causal” categories.
Outcome: The proposed model can be used to identify the inappropriate use of causal language in scientific publications and news articles.
Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals? (2024.findings-acl)

Copied to clipboard

Challenge: Recent work on the evaluation of large language models (LLMs) has shown unprecedented performance on diverse language generation tasks.
Approach: They investigate the controllability of large language models on scientific summarization tasks by controlling stylistic and content coverage factors.
Outcome: The proposed model outperforms humans on the MuP review generation task in terms of similarity to reference summaries and human preferences.
Multi Graph Neural Network for Extractive Long Document Summarization (2022.coling-1)

Copied to clipboard

Challenge: Heterogeneous Graph Neural Networks (GNN) have been proposed as an emergent approach for extracting document summarization (EDS) but there are still limitations in applying it for long documents due to the lack of inter-sentence connections.
Approach: They propose to build a graph on sentence-level nodes and combine it with HeterGNN to capture the semantic information in terms of both inter and intra-sentence connections.
Outcome: Experiments on two datasets show that the proposed method achieves state-of-the-art in this research field.
To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering (2024.acl-long)

Copied to clipboard

Challenge: Medical open-domain question answering requires substantial access to specialized knowledge.
Approach: They propose a framework that generates multiple-choice questions from a set of open-book parameters and a small-scale reader that can outcompete closed-book questions by 706x using fewer parameters.
Outcome: The proposed framework outperforms closed-book models on MedQA-USMLE, MedMCQA, and MMLU while using up to 706x fewer parameters.
LinkBERT: Pretraining Language Models with Document Links (2022.acl-long)

Copied to clipboard

Challenge: Existing language model pretraining methods do not capture dependencies or knowledge that span across documents.
Approach: They propose a language model pretraining method that leverages links between documents . they use masked language modeling and document relation prediction to model LMs .
Outcome: The proposed method outperforms existing methods on downstream tasks across two domains.
An Efficient Coarse-to-Fine Facet-Aware Unsupervised Summarization Framework Based on Semantic Blocks (2022.coling-1)

Copied to clipboard

Challenge: Existing unsupervised summarization methods fail to consider efficiency and effectiveness when the input document is extremely long.
Approach: They propose an efficient Coarse-to-Fine Facet-Aware Ranking framework for unsupervised long document summarization based on the semantic block.
Outcome: The proposed framework can achieve new state-of-the-art unsupervised summarization results on Gov-Report, billSum, arXiv, and PubMed.
Balancing Methods for Multi-label Text Classification with Long-Tailed Class Distribution (2021.emnlp-main)

Copied to clipboard

Challenge: Multi-label text classification is a challenging task because it requires capturing label dependencies.
Approach: They propose to use distribution-balanced loss functions to solve label dependency problems in multi-label text classification by capturing label dependencies from a fixed-set of labels.
Outcome: The proposed loss function addresses both the class imbalance and label linkage problems and outperforms other loss functions.
Profiling Medical Journal Articles Using a Gene Ontology Semantic Tagger (L18-1)

Copied to clipboard

Challenge: a growing number of scientific publications are based on sub-divisions and sub-communities of expertise becoming disconnected from each other.
Approach: They propose to examine corpora derived from bodies of genetics literature and use it to make comparisons and improve retrieval methods.
Outcome: The proposed methods will help to make comparisons and improve retrieval methods using domain knowledge via an existing gene ontology.
Enriching and Controlling Global Semantics for Text Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Abstractive summarization models have been proven effective in creating fluent and informative summaries, but they suffer from the short-range dependency problem, causing them to produce summary that miss the key points of document.
Approach: They propose a neural topic model empowered with normalizing flow to capture global semantics of the document and integrate them into the summarization model.
Outcome: The proposed model outperforms state-of-the-art summarization models on five common text summarizing datasets, namely CNN/DailyMail, XSum, Reddit TIFU, arXiv, and PubMed.
Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modelling (2025.emnlp-main)

Copied to clipboard

Challenge: integrating coreference and decomposition increases recall on rare relations by over 20%.
Approach: They propose an open-source pipeline for extracting sentence-level knowledge graphs by combining robust coreference resolution with syntactic sentence decomposition.
Outcome: The proposed pipeline achieves a 99.8% exact-match accuracy on sentence simplification.
Medical Vision-Language Pre-Training for Brain Abnormalities (2024.lrec-main)

Copied to clipboard

Challenge: Existing vision-language models lack expertise for medical applications due to the scarcity and complexity of data.
Approach: They propose a pipeline to collect medical image-text aligned data for pretraining from public resources such as PubMed and build a high-performance vision-language model tailored to specific medical tasks.
Outcome: The proposed model is based on a large brain image-text dataset and will be released to the public.
PEaCE: A Chemistry-Oriented Dataset for Optical Character Recognition on Scientific Documents (2024.lrec-main)

Copied to clipboard

Challenge: Existing open-source OCR models focus on scientific texts or generic printed English . Nougat is unable to parse tables in PubMed articles .
Approach: They propose to train OCR models for scientific or generic printed English . Nougat is a popular tool for parsing academic documents, but unable to parse PubMed tables .
Outcome: The proposed models perform better when trained on real-world records than those trained on synthetic records.
"Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations (2026.acl-long)

Copied to clipboard

Challenge: Existing large language models (LLMs) are reactive and respond only when prompted, limiting their effectiveness in collaborative settings.
Approach: They introduce a proactive LLM assistant designed to enhance biomedical collaboration between AI systems and human experts through timely, context-aware interventions.
Outcome: The proposed model outperforms baselines in intervention precision and collaborative task utility, highlighting the potential of proactive LLMs as intelligent scientific assistants.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations