Evaluating Multi-Hop Reasoning in Large Language Models: A Chemistry-Centric Benchmark (2026.findings-eacl)
Copied to clipboard
Mohammad Khodadad, Ali Shiraee Kasmaee, Mahdi Astaraki, Nicholas Sherck, Hamidreza Mahyar, Soheila Samiee
| Challenge: | a new pipeline for compositional multi-hop reasoning in large language models is being developed . a recent study shows that even state-of-the-art models struggle with compositional reasoning . |
| Approach: | They propose a pipeline that builds benchmarks from proprietary or public data . they use generative reasoning models, chemical named-entity recognition, and external knowledge bases to build knowledge graphs. |
| Outcome: | The proposed pipeline compares state-of-the-art models with and without retrieval augmentation . the pipeline is generalizable with fine-tuning, enabling creation of challenging benchmarks . |
Similar Papers
ChemReason-Bench: Benchmarking Large Language Models for Procedural Reasoning in Experimental Chemistry (2026.acl-long)
Copied to clipboard
| Challenge: | Experimental protocols in organic synthesis specify not only the intended transformation, but also an executable sequence of operations and conditions. |
| Approach: | They propose a human-validated benchmark for verifiable experimental procedure reasoning . they instantiate 7306 benchmark tasks across six complementary formats . |
| Outcome: | The proposed benchmarks show that the evaluations are less diagnostic of procedure-level decision making. |
From Generalist to Specialist: A Survey of Large Language Models for Chemistry (2025.coling-main)
Copied to clipboard
| Challenge: | Existing studies on pretraining of LLMs on extensive web-based texts are insufficient for advanced scientific discovery, especially in chemistry. |
| Approach: | They outline methodologies for incorporating domain-specific chemistry knowledge and multi-modal information into LLMs and conceptualize chemistry LLM agents using chemistry tools. |
| Outcome: | The proposed models are based on domain-specific chemistry knowledge and multi-modal information and are capable of accelerating scientific research. |
Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers? (2024.emnlp-main)
Copied to clipboard
| Challenge: | State-of-the-art Large Language Models (LLMs) are accredited with a number of different capabilities, including reading comprehension, mathematical and reasoning skills, and possessing scientific knowledge. |
| Approach: | They propose a benchmark to generate seemingly plausible multi-hop reasoning chains that ultimately lead to incorrect answers. |
| Outcome: | The proposed model circumvents the reasoning requirement but in subtle ways . it shows that it is more difficult to generate plausible alternatives . |
MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Correction (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown growing potential in molecular sciences, but they often produce chemically inaccurate descriptions and struggle to recognize or justify potential errors. |
| Approach: | They propose a benchmark to assess LLMs on error detection and correction in molecular descriptions. |
| Outcome: | The proposed benchmark targets LLMs on error detection and correction in molecular descriptions. |
CheMM-R1: Enhancing Chemical Structure Recognition and Elucidation with Reasoning Multimodal Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing multimodal large language models lack domain-specific expertise to perform chemical tasks. |
| Approach: | They propose a benchmark dataset for evaluating multi-step multimodal reasoning capacities in the chemistry domain. |
| Outcome: | The proposed model surpasses existing models in all CheMM-Bench tasks. |
ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use (2025.acl-long)
Copied to clipboard
Junjie Ye, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zehui Chen, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Siyu Yuan, Tao Gui, Qi Zhang, Xuanjing Huang, Jiecao Chen
| Challenge: | Effective evaluation of multi-hop tool use is critical for analyzing the understanding, reasoning, and function-calling capabilities of large language models. |
| Approach: | They propose a dataset that provides rigorous evaluation of multi-hop tool use. |
| Outcome: | The proposed model achieves 49.04% accuracy across five model families. |
On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert Ratings (2021.emnlp-main)
Copied to clipboard
| Challenge: | a large corpus of domain-expert relevance ratings augments a corpus for compositional explanations . a writer's study shows that the evaluations of compositional inference models underestimate performance . |
| Approach: | They construct a corpus of 126k domain-expert relevance ratings that augment explanations to standardized science exam questions. |
| Outcome: | The results show that evaluations underestimate performance of compositional explanations . they show that models regularly discover and produce valid explanations that are different than gold explanations. |
BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for multi-hop reasoning in biomedical domain are lacking . bioHopR provides benchmarks to evaluate multi-step reasoning in structured biomedic knowledge graphs . |
| Approach: | They propose a benchmark to evaluate multi-hop, multi-answer reasoning in biomedical knowledge graphs. |
| Outcome: | BioHopR evaluates multi-hop reasoning in biomedical knowledge graphs based on the PrimeKG model . it outperforms proprietary models and open-source biomedal models in 1-hop and 2-hop tasks . |
Do Large Language Models Latently Perform Multi-Hop Reasoning? (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of ‘Superstition’ is". |
| Approach: | They examine whether Large Language Models (LLMs) latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of ‘Superstition’ is". |
| Outcome: | The proposed model can latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of ‘Superstition’ is". |
Two Steps from Hell: Compositionality on Chemical LMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Experiments with state-of-the-art ChemLLMs show significant performance drops in compositional tasks, highlighting the need for models that move beyond pattern recognition. |
| Approach: | They introduce a benchmark to evaluate chemical language models' understanding of chemical language by identifying and analyzing compositional patterns within chemical data. |
| Outcome: | The proposed benchmark shows that existing LLMs can handle complex queries without pattern recognition. |