Challenge: Recent advances in LLMs have significantly improved mathematical problem-solving, with models like GPT-4 achieving human-level performance.
Approach: They propose a bilingual English-Korean dataset enriched with teacher solutions, student solutions, and annotations marking students’ initial errors.
Outcome: The proposed model achieves high agreement with human judgments and lower latency and resource usage than commercial APIs, demonstrating strong computational efficiency.

Similar Papers

MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have showcased significant improvements in mathematics, but traditional benchmarks like GSM8k offer a unidimensional perspective.
Approach: MathBench is a benchmark that rigorously assesses the mathematical capabilities of large language models.
Outcome: MathBench spans a wide range of mathematical disciplines, offering a detailed evaluation of both theoretical understanding and practical problem-solving skills.
Exposing the Achilles’ Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluations focus on final accuracy, neglecting the critical aspect of reasoning capabilities.
Approach: They propose to evaluate LLMs’ abilities to detect and correct reasoning mistakes by using rule-based methods and smaller language models.
Outcome: The proposed model outperforms existing models such as GPT-4o and GPT4 in both accuracy and accuracy, but lacks data contamination and memorization concerns.
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks primarily focus on English or a narrow subset of high-resource languages, leaving significant gaps in assessing multilingual and cross-lingual mathematical reasoning.
Approach: They propose a parallel multilingual benchmark for mathematical problem solving and reasoning that encompasses 2,890 parallel Bangla-English gold standard artifacts.
Outcome: The proposed model encompasses 2,890 parallel Bangla-English gold standard artifacts, totaling 30K aligned question–answer pairs across thirteen languages, representing high-, medium-, and low-resource linguistic settings.
Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: The performance of large language models (LLMs) on existing reasoning benchmarks has significantly improved over the past decade.
Approach: They propose a benchmark dataset for evaluating the problem solving abilities of large language models (LLMs) they curate 515 challenging problems from the highly competitive IIT JEE-Advanced exam.
Outcome: The proposed model performs better on open-source and proprietary models than the current model, but with techniques like self-consistency, self-refinement and chain-of-thought prompting.
The Impact of Language on Arithmetic Proficiency: A Multilingual Investigation with Cross-Agent Checking Computation (2024.naacl-short)

Copied to clipboard

Challenge: Large language models (LLMs) have garnered significant attention over the past year . previous studies have evaluated LLMs' performance in solving math word problems, but there is little discussion on whether they comprehend the operations they generate.
Approach: They challenge the notion that arithmetic is language-independent and compare models with cross-agent collaborations to find significant limitations in their performance.
Outcome: The proposed model outperforms collaborative approaches in basic arithmetic tasks.
GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive performance across various mathematical reasoning benchmarks.
Approach: They introduce an adversarial grade school math dataset and explore whether LLMs can be more robust when questions are slightly changed.
Outcome: The proposed method generates and verifies each intermediate thought based on its reasoning goal and calculation result.
CogBench: Benchmarking Cognitive Alignment of Large Language Models in Educational Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) possess strong capabilities in language understanding and generation, as well as remarkable problem-solving abilities.
Approach: They propose a benchmark to assess the cognitive alignment capabilities of large language models in educational QA.
Outcome: The proposed evaluation benchmark assesses the cognitive alignment capabilities of large language models in educational QA.
LLMs cannot spot math errors, even when allowed to peek into the solution (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) demonstrate impressive performance on existing reasoning benchmarks, but struggle with meta-reasoning tasks such as locating the first error step in student solutions.
Approach: They propose an approach that generates an intermediate corrected student solution, aligning more closely with the original student’s solution, which helps improve performance.
Outcome: The proposed approach generates an intermediate corrected student solution, aligning more closely with the original student’s solution, which helps improve performance.
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have demonstrated the potential of large language models (LLMs) for automatic error detection in math word problems (MWPs).
Approach: They propose a framework that generates adaptive reference solutions using LLMs to enhance error detection by reducing conformity bias in MWPs.
Outcome: The proposed framework mitigates the performance gap between conventional and alternative solutions in MWPs, especially when combined with reasoning-enhancing techniques like chain-of-thought prompting.
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) achieve impressive performance on complex benchmarks yet sometimes fail on basic math reasoning.
Approach: They propose a benchmark to evaluate the efficiency of reasoning in large language models . they formalize the accuracy-verbosity tradeoff and introduce the overthinking score .
Outcome: The proposed model performs well on complex benchmarks but fails on basic math reasoning . the proposed model generates 18 more tokens while achieving lower accuracy .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations