Challenge: ConceptMath evaluates concept-wise mathematical reasoning of Large Language Models (LLMs) Existing benchmarks that evaluate general mathematical reasoning with an average accuracy fail to probe the fine-grained failure modes of mathematical reasoning on specific datasets.
Approach: They introduce a bilingual, fine-grained benchmark that evaluates concept-wise mathematical reasoning of Large Language Models.
Outcome: The proposed benchmarks evaluate concept-wise mathematical reasoning of Large Language Models with concept-based accuracies.

Similar Papers

HighMATH: Evaluating Math Reasoning of Large Language Models in Breadth and Depth (2025.findings-emnlp)

Copied to clipboard

Challenge: a gap in math models' accuracy has been widened with the development of large language models (LLMs) . a new study aims to bridge this gap by evaluating a set of high-level math reasoning models .
Approach: They propose to evaluate large language models on existing math benchmarks to bridge this gap . they collect 5,293 problems from Chinese senior high school mathematics exams .
Outcome: The proposed model is based on o1-like models and a high-level model.
TMATH A Dataset for Evaluating Large Language Models in Generating Educational Hints for Math Word Problems (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being applied in education, showing significant potential in personalized instruction, student feedback, and intelligent tutoring systems (ITSs).
Approach: They propose a dataset specifically designed to evaluate LLMs’ ability to generate high-quality hints for Math Word Problems.
Outcome: The proposed dataset shows that LLMs can generate more accurate and contextually appropriate educational hints for math word problems without offering direct answers.
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have showcased significant improvements in mathematics, but traditional benchmarks like GSM8k offer a unidimensional perspective.
Approach: MathBench is a benchmark that rigorously assesses the mathematical capabilities of large language models.
Outcome: MathBench spans a wide range of mathematical disciplines, offering a detailed evaluation of both theoretical understanding and practical problem-solving skills.
ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts (2025.findings-emnlp)

Copied to clipboard

Challenge: ScholarBench evaluates domain-specific knowledge of large language models (LLMs) prior benchmarks lack the scalability to handle complex academic tasks.
Approach: ScholarBench evaluates the academic reasoning ability of large language models . the benchmark is constructed through a three-step process .
Outcome: ScholarBench evaluates the academic reasoning ability of large language models . the benchmark comprises 5,031 examples in Korean and 5,309 examples in English .
HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies show that large language models are robust in commonsense reasoning . however, some variations in questions can lead to incorrect responses .
Approach: They propose a large-scale bilingual benchmark consisting of 11,200 cases . they conduct extensive experiments on 41 representative LLMs .
Outcome: The proposed benchmark systematically evaluates the robustness of large language models in commonsense reasoning.
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks for large reasoning models are saturated by a lack of reliable and verifiable benchmarks.
Approach: They propose a rigorously curated, Olympiad-level math benchmark comprising 350 problems, each with parallel English and Chinese versions.
Outcome: The proposed benchmark unifies two evaluation paradigms and offers 150 problems formalized in Lean 4 for rigorous process-level evaluation.
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks primarily focus on English or a narrow subset of high-resource languages, leaving significant gaps in assessing multilingual and cross-lingual mathematical reasoning.
Approach: They propose a parallel multilingual benchmark for mathematical problem solving and reasoning that encompasses 2,890 parallel Bangla-English gold standard artifacts.
Outcome: The proposed model encompasses 2,890 parallel Bangla-English gold standard artifacts, totaling 30K aligned question–answer pairs across thirteen languages, representing high-, medium-, and low-resource linguistic settings.
MMATH: A Multilingual Benchmark for Mathematical Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: a benchmark for multilingual complex reasoning spans 374 high-quality math problems across 10 typologically diverse languages.
Approach: They propose a benchmark for multilingual complex reasoning across 10 languages . they show reasoning in English and answering in target languages can enhance performance .
Outcome: The proposed benchmark demonstrates that models with high-quality reasoning can perform in multiple languages.
PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Numerical reasoning is ubiquitous in scientific research and financial analysis, but few benchmarks evaluate them by integrating numerical processing and mathematical reasoning.
Approach: They propose a numerically-integrated hierarchical benchmark with 27,215 questions derived from 7,404 math word problems that spans 4 key cognitive aspects, 14 subcategories, and 2 modalities.
Outcome: The proposed model improves Qwen-2.5 score with SOLVE and IRPO training.
UTMath: A Benchmark for Math Evaluation with Unit Test (2025.findings-emnlp)

Copied to clipboard

Challenge: Prevailing benchmarks for mathematical reasoning include MATH and AIME . predicated on single-instantiation problems with fixed numbers, these models leave generalization on isomorphic problem variants untested.
Approach: They propose a mathematical reasoning benchmark that quantifies solution accuracy and solution space generality.
Outcome: The proposed model solves 1,053 problems spanning 9 mathematical domains . the best-performing model solved only 32.57% of the problems .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations