Challenge: Existing multimodal large language models lack domain-specific expertise to perform chemical tasks.
Approach: They propose a benchmark dataset for evaluating multi-step multimodal reasoning capacities in the chemistry domain.
Outcome: The proposed model surpasses existing models in all CheMM-Bench tasks.

Similar Papers

Evaluating Multi-Hop Reasoning in Large Language Models: A Chemistry-Centric Benchmark (2026.findings-eacl)

Copied to clipboard

Challenge: a new pipeline for compositional multi-hop reasoning in large language models is being developed . a recent study shows that even state-of-the-art models struggle with compositional reasoning .
Approach: They propose a pipeline that builds benchmarks from proprietary or public data . they use generative reasoning models, chemical named-entity recognition, and external knowledge bases to build knowledge graphs.
Outcome: The proposed pipeline compares state-of-the-art models with and without retrieval augmentation . the pipeline is generalizable with fine-tuning, enabling creation of challenging benchmarks .
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Currently, vision-Language Models are optimized for direct visual question-answering tasks.
Approach: They propose a visual-language-based VLM that prioritizes reasoning within the perception process.
Outcome: The proposed model outperforms existing models and domain-specific open-source models in the chemical domain.
Boosting LLM’s Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Molecular structure elucidation involves deducing a molecule’s structure from various types of spectral data, which is crucial in chemical experimental analysis.
Approach: They propose a Knowledge-enhanced reasoning framework for Molecular Structure Elucidation that leverages Monte Carlo Tree Search for test-time scaling as a plugin to extend the LLMs’ coverage of the chemical structure space.
Outcome: The proposed framework significantly improves on both GPT-4o-mini and GPT4o, and a specialized molecule-spectrum scorer improves performance.
Structural Reasoning Improves Molecular Understanding of LLM (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown significant performance, approaching human perception levels.
Approach: They propose an approach that sketches molecular structures for reasoning by explicitly incorporating key structural features into the model.
Outcome: The proposed framework improves molecular understanding through extensive experiments.
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Current scientific reasoning models struggle with generalization across domains and fall short of multimodal perception.
Approach: They propose to use multimodal large language models to integrate text, images, and other modalities to enhance scientific reasoning.
Outcome: The proposed models can integrate text, images, and other modalities and improve reasoning across disciplines.
Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) outperform existing benchmarks in both natural language and coding domains.
Approach: They propose a scalable benchmark that integrates vision and language modalities to address this gap by eliminating textual shortcuts.
Outcome: The new benchmark outperforms existing benchmarks in both natural language and coding domains.
ChemReason-Bench: Benchmarking Large Language Models for Procedural Reasoning in Experimental Chemistry (2026.acl-long)

Copied to clipboard

Challenge: Experimental protocols in organic synthesis specify not only the intended transformation, but also an executable sequence of operations and conditions.
Approach: They propose a human-validated benchmark for verifiable experimental procedure reasoning . they instantiate 7306 benchmark tasks across six complementary formats .
Outcome: The proposed benchmarks show that the evaluations are less diagnostic of procedure-level decision making.
MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Correction (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown growing potential in molecular sciences, but they often produce chemically inaccurate descriptions and struggle to recognize or justify potential errors.
Approach: They propose a benchmark to assess LLMs on error detection and correction in molecular descriptions.
Outcome: The proposed benchmark targets LLMs on error detection and correction in molecular descriptions.
Improving Chemical Understanding of LLMs via SMILES Parsing (2025.emnlp-main)

Copied to clipboard

Challenge: Molecular string representations such as SMILES and SELFIES are becoming a standard format for applying large language models (LLMs) however, molecular strings follow complex syntactic rules for encoding molecules, which LLMs struggle to interpret.
Approach: They propose a framework that parses SMILES into clean and deterministic tasks to promote graph-level molecular comprehension.
Outcome: The proposed framework improves structural comprehension and competes with the baseline on the Mol-Instructions benchmark.
REAP: Towards Effective Training-Free Chemical Reasoning with Explicit Atomic Priors (2026.findings-acl)

Copied to clipboard

Challenge: Current approaches to instill explicit priors into LLMs often suffer from an information bottleneck .
Approach: They propose a training-free framework that equips LLMs with an external knowledge base, enabling them to reason over retrieved chemical priors dynamically.
Outcome: Experiments show that REAP outperforms current reasoning methods and rivals state-of-the-art training-based models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations