Challenge: Existing MWP corpora are limited in language patterns and problem types . a new corpus of 2,305 MWps is proposed that is more diverse in terms of lexicon usage .
Approach: They propose to use ASDiv to measure lexicon usage diversity of a given MWP corpus.
Outcome: The proposed corpus covers more problem types and text patterns than existing corpora and reflects the true capability of solvers more faithfully.

Similar Papers

It Ain’t Over: A Multi-aspect Diverse Math Word Problem Dataset (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies lack diversity in problem types, lexical usage patterns, languages, and intermediate solution forms for the math word problem.
Approach: They propose a new MWP dataset with a wide range of diversity in problem types, lexical usage patterns, languages, and intermediate solutions.
Outcome: The proposed dataset provides an opportunity to evaluate the capability of large language models.
Learning by Analogy: Diverse Questions Generation in Math Word Problem (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for solving math word problem (MWP) use shortcut learning to train solvers based on samples with a single question.
Approach: They propose to generate diverse yet consistent questions from a common scenario . they then feed the equations to a question generator to obtain the diverse questions . their method leads to performance improvement on the current benchmark Math23K .
Outcome: The proposed method generates diverse yet consistent questions with a variety of equations and questions . it improves on the current benchmark, which is based on the proposed method .
What Makes Math Word Problems Challenging for LLMs? (2024.findings-naacl)

Copied to clipboard

Challenge: Experiments show that even quite powerful LLMs are still challenged by MWPs.
Approach: They propose to analyze what makes math word problems (MWPs) in English challenging for large language models (LLMs).
Outcome: The proposed model can handle a range of core NLP tasks, but it has emergent abilities, such as ability to solve mathematical puzzles.
Are NLP Models really able to Solve Simple Math Word Problems? (2021.naacl-main)

Copied to clipboard

Challenge: Existing solvers for math word problems often achieve high performance on benchmark datasets . existing models rely on shallow heuristics to achieve high accuracy .
Approach: They restrict their attention to English MWPs taught in grades four and lower . they propose a challenge dataset to test the accuracy of MWp solvers .
Outcome: The proposed model can solve a large fraction of MWPs even with shallow heuristics . the proposed model is much lower on the challenge dataset SVAMP .
Math Word Problem Solving by Generating Linguistic Variants of Problem Statements (2023.acl-srw)

Copied to clipboard

Challenge: Existing models for solving Math Word Problems depend on shallow heuristics and spurious correlations to derive the solution expressions.
Approach: They propose a framework for MWP solvers based on generation of linguistic variants of problem text.
Outcome: The proposed framework improves the mathematical reasoning and robustness of the proposed model.
ArMATH: a Dataset for Solving Arabic Math Word Problems (2022.lrec-1)

Copied to clipboard

Challenge: This paper is the first to use deep learning methods to solve Arabic MWPs . it is also the first study to use transfer learning to solve MWp across different languages .
Approach: They contribute to the first large-scale dataset for Arabic Math Word Problems . they use deep learning methods to solve Arabic MWPs and a transfer learning model to promote performance .
Outcome: The proposed model improves Arabic MWP solvers by 3% over the existing model.
The ACQDIV Corpus Database and Aggregation Pipeline (2020.lrec-1)

Copied to clipboard

Challenge: ACQDIV corpus database and aggregation pipeline aims to identify universal cognitive processes that allow children to acquire any language.
Approach: They present the ACQDIV corpus database and aggregation pipeline . the tool aims to identify universal cognitive processes that allow children to acquire any language .
Outcome: The ACQDIV corpus database and aggregation pipeline is a tool developed by the European Research Council . the database represents 15 corpora from 14 typologically maximally diverse languages .
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
TMATH A Dataset for Evaluating Large Language Models in Generating Educational Hints for Math Word Problems (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being applied in education, showing significant potential in personalized instruction, student feedback, and intelligent tutoring systems (ITSs).
Approach: They propose a dataset specifically designed to evaluate LLMs’ ability to generate high-quality hints for Math Word Problems.
Outcome: The proposed dataset shows that LLMs can generate more accurate and contextually appropriate educational hints for math word problems without offering direct answers.
EDUMATH: Generating Standards-aligned Educational Math Word Problems (2026.acl-long)

Copied to clipboard

Challenge: Math word problems (MWPs) are critical elements of K-12 math education and can be customized to students' interests and ability levels.
Approach: They propose that LLMs can generate MWPs customized to student interests and math education standards by using an open and closed LLM to evaluate over 11,000 MWps and develop a teacher-annotated dataset for standards-aligned educational MWPS generation.
Outcome: The proposed model outperforms existing closed models without training and is more similar to human-written MWPs but prefers customized MWPS with grade school students.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations