Challenge: Advances in large language models have spurred research into enhancing their reasoning capabilities, particularly in math-rich STEM documents.
Approach: They propose a benchmark dataset to evaluate LLMs’ reasoning abilities on math symbols within contextual scientific text.
Outcome: The proposed dataset demonstrates that state-of-the-art LLMs achieve an average accuracy of 20-60% under in-context learning and 50-60% with fine-tuning, highlighting a substantial gap in their ability to classify mathematical symbols.

Similar Papers

LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated strong performance across various natural language processing tasks, but their proficiency in mathematical reasoning remains a key challenge.
Approach: They propose a process-oriented framework to evaluate LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth.
Outcome: The proposed framework evaluates LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth.
SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has addressed problems in unstructured grounding, multi-equation dependency, and human-aligned evaluation.
Approach: They construct a dataset of scientific texts and evaluate it using an explainable equation generation workflow using automatic metrics and human judgments.
Outcome: The proposed model achieves moderate performance on lexical and syntactic similarity, but struggles with semantic accuracy.
Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have limitations when it comes to comprehending and expressing world knowledge that extends beyond the boundaries of natural language.
Approach: They propose a model that integrates symbolic data into LLM training without loss of generality ability.
Outcome: The proposed model performs better on symbol- and NL-centric tasks.
MARIO: MAth Reasoning with code Interpreter Output - A Reproducible Pipeline (2024.findings-acl)

Copied to clipboard

Challenge: Large language models lack mathematical reasoning, a hurdle on the path to true artificial general intelligence.
Approach: They propose a protocol for fine-tuning large language models with a Python code interpreter to enhance the text analysis of the LLMs.
Outcome: The proposed protocol improves the performance of a 7B-parameter LLM on the GSM8K and MATH datasets while allowing for an outlier-free value model-based inference method.
Leveraging Language-based Representations for Better Solving Symbol-related Problems with Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Symbols are used in abstract reasoning, chemical property prediction, and tabular question-answering.
Approach: They propose a method that converts symbols to language-based representations to improve their accuracy.
Outcome: The proposed method improves the accuracy of symbols in language-based models.
Disentangling Text and Math in Word Problems: Evidence for the Bidimensional Structure of Large Language Models’ Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies show that LLMs struggle with text interpretation and equation solving, despite distinct proficiencies in textual and mathematical components.
Approach: They disentangle textual interpretation and mathematical solving steps in word problems drawn from Brazil's largest college entrance exam and popular grade school-level benchmark GSM8K.
Outcome: The proposed model outperforms LLMs in Brazil's largest college entrance exam and popular grade school-level benchmark.
MathFish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula (2024.findings-emnlp)

Copied to clipboard

Challenge: pedagogical experts spend months reviewing published math problems to ensure that they align with critical skills or concepts.
Approach: They propose a novel approach for evaluating language models' mathematical abilities by combining a dataset of 385 fine-grained descriptions of K-12 math skills and concepts with 9.9K math problems labeled with these standards.
Outcome: The proposed model can discern skills and concepts enabled by math content, and it can be used to assess language models' mathematical abilities.
LLM Parameters for Math Across Languages: Shared or Separate? (2026.acl-srw)

Copied to clipboard

Challenge: Existing research on large language models (LLMs) has focused on performance or representational properties, but it remains unclear whether these differences reflect language-specific parameters or a shared mechanism.
Approach: They propose to localize and compare model parameters that support mathematical reasoning across languages.
Outcome: The proposed analysis shows that the model parameters in English and lower-resource languages exhibit partial cross-lingual overlap with systematic language-dependent differences.
NEWTON: Are Large Language Models Capable of Physical Reasoning? (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have been shown to encapsulate syntactic, semantic, word sense, and common-sense knowledge, but limited exploration of their physical reasoning abilities has been conducted.
Approach: They propose a repository and benchmark to evaluate LLMs' physical reasoning skills . they use a pipeline to generate a variant of the benchmark customized to the objects and attributes relevant for their application.
Outcome: The proposed benchmark examines the reasoning capabilities of language models across reasoning tasks.
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) achieve impressive performance on complex benchmarks yet sometimes fail on basic math reasoning.
Approach: They propose a benchmark to evaluate the efficiency of reasoning in large language models . they formalize the accuracy-verbosity tradeoff and introduce the overthinking score .
Outcome: The proposed model performs well on complex benchmarks but fails on basic math reasoning . the proposed model generates 18 more tokens while achieving lower accuracy .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations