What Makes Math Word Problems Challenging for LLMs? (2024.findings-naacl)

Copied to clipboard

Challenge: Experiments show that even quite powerful LLMs are still challenged by MWPs.
Approach: They propose to analyze what makes math word problems (MWPs) in English challenging for large language models (LLMs).
Outcome: The proposed model can handle a range of core NLP tasks, but it has emergent abilities, such as ability to solve mathematical puzzles.

Similar Papers

Are NLP Models really able to Solve Simple Math Word Problems? (2021.naacl-main)

Copied to clipboard

Challenge: Existing solvers for math word problems often achieve high performance on benchmark datasets . existing models rely on shallow heuristics to achieve high accuracy .
Approach: They restrict their attention to English MWPs taught in grades four and lower . they propose a challenge dataset to test the accuracy of MWp solvers .
Outcome: The proposed model can solve a large fraction of MWPs even with shallow heuristics . the proposed model is much lower on the challenge dataset SVAMP .
Large Language Models for Mathematical Reasoning: Progresses and Challenges (2024.eacl-srw)

Copied to clipboard

Challenge: a survey examines the landscape of mathematical problem-solving techniques . large language models have proven to be potent assets in unraveling nuances of mathematical reasoning .
Approach: They examine the evolution of Large Language Models (LLMs) for solving mathematical problems . they examine the spectrum of LLM-oriented techniques proposed for solving math problems - and their challenges .
Outcome: The survey examines the spectrum of proposed LLM-oriented techniques in solving math problems.
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework (2025.findings-emnlp)

Copied to clipboard

Challenge: Current error classification methods rely on static and predefined categories to capture error patterns.
Approach: They propose a framework for automated dynamic error classification in mathematical reasoning that incorporates common error patterns as explicit guidance.
Outcome: The proposed framework reduces human bias and fine-grained analysis of error patterns.
Three Questions Concerning the Use of Large Language Models to Facilitate Mathematics Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: After the pandemic, e-learning has become part of mainstream education.
Approach: They propose to integrate large language models (LLMs) into educational settings to enhance students' mathematical problem-solving skills by providing adaptive feedback.
Outcome: The proposed model can generate free-text rationalizations and misinterpret meanings and can also misinterprét students' answers.
Exposing the Achilles’ Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluations focus on final accuracy, neglecting the critical aspect of reasoning capabilities.
Approach: They propose to evaluate LLMs’ abilities to detect and correct reasoning mistakes by using rule-based methods and smaller language models.
Outcome: The proposed model outperforms existing models such as GPT-4o and GPT4 in both accuracy and accuracy, but lacks data contamination and memorization concerns.
What Kind of Language Is Hard to Language-Model? (P19-1)

Copied to clipboard

Challenge: a recent study suggests that language models perform poorly across languages.
Approach: They propose a model that fits a paired-sample multiplicative mixed-effects model to obtain language difficulty coefficients from at least-pairwise parallel corpora.
Outcome: The proposed model is able to handle missing data and is aware of inter-sentence variation.
EDUMATH: Generating Standards-aligned Educational Math Word Problems (2026.acl-long)

Copied to clipboard

Challenge: Math word problems (MWPs) are critical elements of K-12 math education and can be customized to students' interests and ability levels.
Approach: They propose that LLMs can generate MWPs customized to student interests and math education standards by using an open and closed LLM to evaluate over 11,000 MWps and develop a teacher-annotated dataset for standards-aligned educational MWPS generation.
Outcome: The proposed model outperforms existing closed models without training and is more similar to human-written MWPs but prefers customized MWPS with grade school students.
Rationales for Answers to Simple Math Word Problems Confuse Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies show that large language models have advanced mathematical problem-solving abilities in grade school math word problems.
Approach: They propose to combine fine-tuning and prompt-based methods to improve performance . they propose to use a hybrid algorithm to fine- tune LLMs on specific tasks .
Outcome: The proposed methods improve performance on the proposed reasoning process evaluation benchmarks.
TMATH A Dataset for Evaluating Large Language Models in Generating Educational Hints for Math Word Problems (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being applied in education, showing significant potential in personalized instruction, student feedback, and intelligent tutoring systems (ITSs).
Approach: They propose a dataset specifically designed to evaluate LLMs’ ability to generate high-quality hints for Math Word Problems.
Outcome: The proposed dataset shows that LLMs can generate more accurate and contextually appropriate educational hints for math word problems without offering direct answers.
Disentangling Text and Math in Word Problems: Evidence for the Bidimensional Structure of Large Language Models’ Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies show that LLMs struggle with text interpretation and equation solving, despite distinct proficiencies in textual and mathematical components.
Approach: They disentangle textual interpretation and mathematical solving steps in word problems drawn from Brazil's largest college entrance exam and popular grade school-level benchmark GSM8K.
Outcome: The proposed model outperforms LLMs in Brazil's largest college entrance exam and popular grade school-level benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations