Assay2Mol: Large Language Model-based Drug Design Using BioAssay Context (2025.emnlp-main)
Copied to clipboard
| Challenge: | Unstructured text that describes biological mechanisms of assays is untapped for early-stage drug discovery. |
| Approach: | They propose a large language model-based workflow that can capitalize on existing biochemical screening assays for early-stage drug discovery. |
| Outcome: | Assay2Mol outperforms machine learning approaches that generate candidate compounds for protein structures while promoting more synthesizable molecule generation. |
Similar Papers
InstructMol: Multi-Modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery (2025.coling-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can attain professional-level proficiency in specific domains through fine-tuning. |
| Approach: | They propose a multi-modal LLM that aligns molecular structures with natural language via an instruction-tuning approach. |
| Outcome: | InstructMol surpasses existing models and reduces the gap with specialists in drug discovery tasks. |
Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing databases contain tens of millions of molecules; PubChem alone has 110 million compounds. |
| Approach: | They propose a task to retrieve molecules using natural language descriptions as queries . they construct a paired dataset of molecules and their corresponding text descriptions . |
| Outcome: | The proposed approach improves results from 0.372 to 0.499 MRR. |
CAGenMol: Condition-Aware Diffusion Language Model for Goal-Directed Molecular Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to optimize target-directed molecular generation fail to reconcile conflicting objectives without compromising structural validity. |
| Approach: | They propose a condition-aware discrete diffusion framework that allows for conditional denoising guided by heterogeneous structural and property signals. |
| Outcome: | The proposed framework improves on structure-conditioned, property-conditioned and dual-conditioned benchmarks in binding affinity, drug-likeness, and success rate. |
A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to Optimization (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are introducing a paradigm shift in molecular discovery by enabling text-guided interaction with chemical spaces through natural language and symbolic notations. |
| Approach: | They analyze the current LLM learning paradigms to tackle four critical evaluation dimensions that have emerged as critical dimensions in recent studies. |
| Outcome: | The proposed models are able to interact with chemical spaces through natural language and symbolic notations, and have emerging extensions to incorporate multi-modal inputs. |
From Knowledge to Treatment: Large Language Model Assisted Biomedical Concept Representation for Drug Repurposing (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for drug repurposing ignore common-sense biomedical concept knowledge in real-world labs, such as mechanistic priors indicating that certain drugs are fundamentally incompatible with specific treatments. |
| Approach: | They propose a Large Language Model-assisted framework for Drug Repurposing which improves the representation of biomedical concepts within KGs. |
| Outcome: | The proposed framework improves the representation of biomedical concepts within KGs by extracting treatment-related textual representations of biomedic entities from large language models and fine-tuning knowledge graph embedding models. |
Language + Molecules (2024.eacl-tutorials)
Copied to clipboard
| Challenge: | In the last year, instruction-following language models have surged in popularity. |
| Approach: | This tutorial will provide an introduction to applying natural language-driven solutions to chemistry problems. |
| Outcome: | This tutorial will provide an introduction to this area of research. it requires no knowledge outside mainstream NLP, and it will enable participants to begin exploring relevant research. |
BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | despite the success of large language models, their performance in highly specialized domains remains unsatisfactory. |
| Approach: | They propose a biomedical tool-calling dataset designed for fine-tuning LLMs . the dataset contains 34 frequently used tools from the NCBI, Ensembl, and UniProt databases . |
| Outcome: | The proposed dataset outperforms commercial LLMs on biomedical domains. |
From Generalist to Specialist: A Survey of Large Language Models for Chemistry (2025.coling-main)
Copied to clipboard
| Challenge: | Existing studies on pretraining of LLMs on extensive web-based texts are insufficient for advanced scientific discovery, especially in chemistry. |
| Approach: | They outline methodologies for incorporating domain-specific chemistry knowledge and multi-modal information into LLMs and conceptualize chemistry LLM agents using chemistry tools. |
| Outcome: | The proposed models are based on domain-specific chemistry knowledge and multi-modal information and are capable of accelerating scientific research. |
LIDDIA: Language-based Intelligent Drug Discovery Agent (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in artificial intelligence for chemistry have sought to expedite individual drug discovery tasks. |
| Approach: | They propose an autonomous agent capable of intelligently navigating the drug discovery process in silico. |
| Outcome: | The proposed agent can generate molecules meeting key pharmaceutical criteria on over 70% of 30 clinically relevant targets and intelligently balances exploration and exploitation in the chemical space. |
Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable Text (2025.naacl-industry)
Copied to clipboard
Ala Jararweh, Oladimeji Macaulay, David Arredondo, Yue Hu, Luis E Tafoya, Kushal Virupakshappa, Avinash Sahu
| Challenge: | Proteins play critical roles in biological systems, yet 99.7% of 227 million known protein sequences remain uncharacterized due to the limitations of experimental methods. |
| Approach: | They propose a multimodal large language model that interprets protein sequences and generates informative text to address open-ended questions about protein functions and attributes. |
| Outcome: | The proposed model outperforms existing models in open-ended question-answering tasks. |