Challenge: Unstructured text that describes biological mechanisms of assays is untapped for early-stage drug discovery.
Approach: They propose a large language model-based workflow that can capitalize on existing biochemical screening assays for early-stage drug discovery.
Outcome: Assay2Mol outperforms machine learning approaches that generate candidate compounds for protein structures while promoting more synthesizable molecule generation.

Similar Papers

InstructMol: Multi-Modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) can attain professional-level proficiency in specific domains through fine-tuning.
Approach: They propose a multi-modal LLM that aligns molecular structures with natural language via an instruction-tuning approach.
Outcome: InstructMol surpasses existing models and reduces the gap with specialists in drug discovery tasks.
Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries (2021.emnlp-main)

Copied to clipboard

Challenge: Existing databases contain tens of millions of molecules; PubChem alone has 110 million compounds.
Approach: They propose a task to retrieve molecules using natural language descriptions as queries . they construct a paired dataset of molecules and their corresponding text descriptions .
Outcome: The proposed approach improves results from 0.372 to 0.499 MRR.
CAGenMol: Condition-Aware Diffusion Language Model for Goal-Directed Molecular Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to optimize target-directed molecular generation fail to reconcile conflicting objectives without compromising structural validity.
Approach: They propose a condition-aware discrete diffusion framework that allows for conditional denoising guided by heterogeneous structural and property signals.
Outcome: The proposed framework improves on structure-conditioned, property-conditioned and dual-conditioned benchmarks in binding affinity, drug-likeness, and success rate.
A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to Optimization (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are introducing a paradigm shift in molecular discovery by enabling text-guided interaction with chemical spaces through natural language and symbolic notations.
Approach: They analyze the current LLM learning paradigms to tackle four critical evaluation dimensions that have emerged as critical dimensions in recent studies.
Outcome: The proposed models are able to interact with chemical spaces through natural language and symbolic notations, and have emerging extensions to incorporate multi-modal inputs.
From Knowledge to Treatment: Large Language Model Assisted Biomedical Concept Representation for Drug Repurposing (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for drug repurposing ignore common-sense biomedical concept knowledge in real-world labs, such as mechanistic priors indicating that certain drugs are fundamentally incompatible with specific treatments.
Approach: They propose a Large Language Model-assisted framework for Drug Repurposing which improves the representation of biomedical concepts within KGs.
Outcome: The proposed framework improves the representation of biomedical concepts within KGs by extracting treatment-related textual representations of biomedic entities from large language models and fine-tuning knowledge graph embedding models.
Language + Molecules (2024.eacl-tutorials)

Copied to clipboard

Challenge: In the last year, instruction-following language models have surged in popularity.
Approach: This tutorial will provide an introduction to applying natural language-driven solutions to chemistry problems.
Outcome: This tutorial will provide an introduction to this area of research. it requires no knowledge outside mainstream NLP, and it will enable participants to begin exploring relevant research.
BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: despite the success of large language models, their performance in highly specialized domains remains unsatisfactory.
Approach: They propose a biomedical tool-calling dataset designed for fine-tuning LLMs . the dataset contains 34 frequently used tools from the NCBI, Ensembl, and UniProt databases .
Outcome: The proposed dataset outperforms commercial LLMs on biomedical domains.
From Generalist to Specialist: A Survey of Large Language Models for Chemistry (2025.coling-main)

Copied to clipboard

Challenge: Existing studies on pretraining of LLMs on extensive web-based texts are insufficient for advanced scientific discovery, especially in chemistry.
Approach: They outline methodologies for incorporating domain-specific chemistry knowledge and multi-modal information into LLMs and conceptualize chemistry LLM agents using chemistry tools.
Outcome: The proposed models are based on domain-specific chemistry knowledge and multi-modal information and are capable of accelerating scientific research.
LIDDIA: Language-based Intelligent Drug Discovery Agent (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in artificial intelligence for chemistry have sought to expedite individual drug discovery tasks.
Approach: They propose an autonomous agent capable of intelligently navigating the drug discovery process in silico.
Outcome: The proposed agent can generate molecules meeting key pharmaceutical criteria on over 70% of 30 clinically relevant targets and intelligently balances exploration and exploitation in the chemical space.
Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable Text (2025.naacl-industry)

Copied to clipboard

Challenge: Proteins play critical roles in biological systems, yet 99.7% of 227 million known protein sequences remain uncharacterized due to the limitations of experimental methods.
Approach: They propose a multimodal large language model that interprets protein sequences and generates informative text to address open-ended questions about protein functions and attributes.
Outcome: The proposed model outperforms existing models in open-ended question-answering tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations