Challenge: Molecular property prediction plays a crucial role in medicinal chemistry . traditional machine learning approaches do not involve natural language .
Approach: They compare performance of four state-of-the-art LLMs on molecular property prediction tasks . they find statistically significant zero- and few-shot preferences for InChI and IUPAC names .
Outcome: The proposed model outperforms the current model on molecular property prediction tasks . the model's representation preferences are based on representation granularity, tokenization and prevalence in pretraining corpora .

Similar Papers

A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to Optimization (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are introducing a paradigm shift in molecular discovery by enabling text-guided interaction with chemical spaces through natural language and symbolic notations.
Approach: They analyze the current LLM learning paradigms to tackle four critical evaluation dimensions that have emerged as critical dimensions in recent studies.
Outcome: The proposed models are able to interact with chemical spaces through natural language and symbolic notations, and have emerging extensions to incorporate multi-modal inputs.
\mathtt{GeLLM^3O}: Generalizing Large Language Models for Multi-property Molecule Optimization (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have remarkable out-of-domain generalizability to novel optimization tasks.
Approach: They propose a series of instruction-tuned LLMs for molecule optimization that outperform state-of-the-art instruction-based LLM models.
Outcome: mathttMuMOInstruct outperforms state-of-the-art LLMs on 5 in-domain and 5 out-of domain tasks.
Predicting Language Models’ Success at Zero-Shot Probabilistic Prediction (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent work has investigated the capabilities of large language models (LLMs) as zero-shot models for generating individual-level characteristics.
Approach: They conduct a large-scale empirical study of large language models’ zero-shot predictive capabilities across a wide range of tabular prediction tasks.
Outcome: The results show that LLMs perform well on the base prediction task, and when they perform well, they are more likely to provide high-quality predictions.
Large Language Models for Controllable Multi-property Multi-objective Molecule Optimization (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for molecule optimization fail to capture property-specific objectives . a series of instruction-tuned LLMs can perform targeted property-specific optimization .
Approach: They propose a set of instruction-tuned LLMs that can perform targeted property-specific optimization.
Outcome: a new instruction-tuned LLM can perform targeted property-specific optimization.
Improving Chemical Understanding of LLMs via SMILES Parsing (2025.emnlp-main)

Copied to clipboard

Challenge: Molecular string representations such as SMILES and SELFIES are becoming a standard format for applying large language models (LLMs) however, molecular strings follow complex syntactic rules for encoding molecules, which LLMs struggle to interpret.
Approach: They propose a framework that parses SMILES into clean and deterministic tasks to promote graph-level molecular comprehension.
Outcome: The proposed framework improves structural comprehension and competes with the baseline on the Mol-Instructions benchmark.
Moleco: Molecular Contrastive Learning with Chemical Language Models for Molecular Property Prediction (2024.emnlp-industry)

Copied to clipboard

Challenge: Pre-trained chemical language models (CLMs) are highly effective in predicting molecular properties, but implicitly contain limited structural information.
Approach: They propose a contrastive learning framework to enhance the understanding of molecular structures within chemical language models by measuring the similarity of fingerprint vectors among different molecules.
Outcome: The proposed framework outperforms state-of-the-art models in molecular property prediction.
Dissecting Human and LLM Preferences (2024.acl-long)

Copied to clipboard

Challenge: a recent study shows that human and Large Language Model preferences are important for model fine-tuning and evaluation.
Approach: They dissect the preferences of human and 32 different Large Language Models to understand their quantitative composition.
Outcome: The proposed model is compared with 32 different large language models using real-world user-model conversations.
How to Make Large Language Models Generate 100% Valid Molecules? (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) can learn to perform a wide range of tasks, but generating valid molecules using representations like SMILES is challenging in few-shot settings.
Approach: They propose a language framework that converts invalid SMILES to SELFIES and LLMs as post-hoc correctors to ensure that the molecules generated by LLM are 100% valid.
Outcome: The proposed model performs worse with SELFIES than with SMILES and improves on other metrics.
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation (2025.findings-emnlp)

Copied to clipboard

Challenge: generative large language models are increasingly used for data augmentation tasks . text samples are mostly selected randomly and a comprehensive overview of other sample selection strategies is lacking.
Approach: They compare random sample selection strategies and random sample sampling strategies to evaluate their effects in a low-resource setting.
Outcome: The proposed model performance improvements are compared with other sample selection strategies.
Characterizing Human and Zero-Shot GPT-3.5 Object-Similarity Judgments (2024.findings-naacl)

Copied to clipboard

Challenge: Recent advances in large language models have yielded few-shot, human-comparable performance on a range of tasks, but studies of LLM annotation accuracy and behavior are sparse.
Approach: They characterize OpenAI’s GPT-3.5’s judgment on a behavioral task for implicit object categorization and give similarities and differences between them.
Outcome: The proposed model augments human responses with LLMs for domains where data is sparse or compute resources are low.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations