STEM-POM: Evaluating Language Models Math-Symbol Reasoning in Document Parsing (2025.findings-acl)
Copied to clipboard
| Challenge: | Advances in large language models have spurred research into enhancing their reasoning capabilities, particularly in math-rich STEM documents. |
| Approach: | They propose a benchmark dataset to evaluate LLMs’ reasoning abilities on math symbols within contextual scientific text. |
| Outcome: | The proposed dataset demonstrates that state-of-the-art LLMs achieve an average accuracy of 20-60% under in-context learning and 50-60% with fine-tuning, highlighting a substantial gap in their ability to classify mathematical symbols. |
Similar Papers
LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated strong performance across various natural language processing tasks, but their proficiency in mathematical reasoning remains a key challenge. |
| Approach: | They propose a process-oriented framework to evaluate LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth. |
| Outcome: | The proposed framework evaluates LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth. |
SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work has addressed problems in unstructured grounding, multi-equation dependency, and human-aligned evaluation. |
| Approach: | They construct a dataset of scientific texts and evaluate it using an explainable equation generation workflow using automatic metrics and human judgments. |
| Outcome: | The proposed model achieves moderate performance on lexical and syntactic similarity, but struggles with semantic accuracy. |
Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have limitations when it comes to comprehending and expressing world knowledge that extends beyond the boundaries of natural language. |
| Approach: | They propose a model that integrates symbolic data into LLM training without loss of generality ability. |
| Outcome: | The proposed model performs better on symbol- and NL-centric tasks. |
MARIO: MAth Reasoning with code Interpreter Output - A Reproducible Pipeline (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models lack mathematical reasoning, a hurdle on the path to true artificial general intelligence. |
| Approach: | They propose a protocol for fine-tuning large language models with a Python code interpreter to enhance the text analysis of the LLMs. |
| Outcome: | The proposed protocol improves the performance of a 7B-parameter LLM on the GSM8K and MATH datasets while allowing for an outlier-free value model-based inference method. |
Leveraging Language-based Representations for Better Solving Symbol-related Problems with Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Symbols are used in abstract reasoning, chemical property prediction, and tabular question-answering. |
| Approach: | They propose a method that converts symbols to language-based representations to improve their accuracy. |
| Outcome: | The proposed method improves the accuracy of symbols in language-based models. |
Disentangling Text and Math in Word Problems: Evidence for the Bidimensional Structure of Large Language Models’ Reasoning (2025.findings-acl)
Copied to clipboard
Pedro Calais, Gabriel Franco, Zilu Tang, Themistoklis Nikas, Wagner Meira Jr., Evimaria Terzi, Mark Crovella
| Challenge: | Existing studies show that LLMs struggle with text interpretation and equation solving, despite distinct proficiencies in textual and mathematical components. |
| Approach: | They disentangle textual interpretation and mathematical solving steps in word problems drawn from Brazil's largest college entrance exam and popular grade school-level benchmark GSM8K. |
| Outcome: | The proposed model outperforms LLMs in Brazil's largest college entrance exam and popular grade school-level benchmark. |
MathFish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula (2024.findings-emnlp)
Copied to clipboard
| Challenge: | pedagogical experts spend months reviewing published math problems to ensure that they align with critical skills or concepts. |
| Approach: | They propose a novel approach for evaluating language models' mathematical abilities by combining a dataset of 385 fine-grained descriptions of K-12 math skills and concepts with 9.9K math problems labeled with these standards. |
| Outcome: | The proposed model can discern skills and concepts enabled by math content, and it can be used to assess language models' mathematical abilities. |
LLM Parameters for Math Across Languages: Shared or Separate? (2026.acl-srw)
Copied to clipboard
Behzad Shomali, Luisa Victor, Tim Selbach, Ali Hamza Bashir, David Berghaus, Joachim Koehler, Mehdi Ali, Markus Frey
| Challenge: | Existing research on large language models (LLMs) has focused on performance or representational properties, but it remains unclear whether these differences reflect language-specific parameters or a shared mechanism. |
| Approach: | They propose to localize and compare model parameters that support mathematical reasoning across languages. |
| Outcome: | The proposed analysis shows that the model parameters in English and lower-resource languages exhibit partial cross-lingual overlap with systematic language-dependent differences. |
NEWTON: Are Large Language Models Capable of Physical Reasoning? (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models have been shown to encapsulate syntactic, semantic, word sense, and common-sense knowledge, but limited exploration of their physical reasoning abilities has been conducted. |
| Approach: | They propose a repository and benchmark to evaluate LLMs' physical reasoning skills . they use a pipeline to generate a variant of the benchmark customized to the objects and attributes relevant for their application. |
| Outcome: | The proposed benchmark examines the reasoning capabilities of language models across reasoning tasks. |
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) achieve impressive performance on complex benchmarks yet sometimes fail on basic math reasoning. |
| Approach: | They propose a benchmark to evaluate the efficiency of reasoning in large language models . they formalize the accuracy-verbosity tradeoff and introduce the overthinking score . |
| Outcome: | The proposed model performs well on complex benchmarks but fails on basic math reasoning . the proposed model generates 18 more tokens while achieving lower accuracy . |