Challenge: Recent advances in natural language processing (NLP) can be attributed to massive scaling of Large Language Models (LLMs).
Approach: They propose a technique that improves performance of Large Language Models (LLMs) on arithmetic problems along with increased reliance in the predictions.
Outcome: The proposed technique improves performance on arithmetic problems and increases confidence in the output results.

Similar Papers

GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive performance across various mathematical reasoning benchmarks.
Approach: They introduce an adversarial grade school math dataset and explore whether LLMs can be more robust when questions are slightly changed.
Outcome: The proposed method generates and verifies each intermediate thought based on its reasoning goal and calculation result.
Large Language Models are Better Reasoners with Self-Verification (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to solve complex natural language processing tasks require multiple steps to verify the answers.
Approach: They propose to use chain of thought prompting to solve reasoning tasks with large language models.
Outcome: The proposed method can improve reasoning performance on arithmetic, commonsense, and logical reasoning datasets.
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) achieve impressive performance on complex benchmarks yet sometimes fail on basic math reasoning.
Approach: They propose a benchmark to evaluate the efficiency of reasoning in large language models . they formalize the accuracy-verbosity tradeoff and introduce the overthinking score .
Outcome: The proposed model performs well on complex benchmarks but fails on basic math reasoning . the proposed model generates 18 more tokens while achieving lower accuracy .
MARIO: MAth Reasoning with code Interpreter Output - A Reproducible Pipeline (2024.findings-acl)

Copied to clipboard

Challenge: Large language models lack mathematical reasoning, a hurdle on the path to true artificial general intelligence.
Approach: They propose a protocol for fine-tuning large language models with a Python code interpreter to enhance the text analysis of the LLMs.
Outcome: The proposed protocol improves the performance of a 7B-parameter LLM on the GSM8K and MATH datasets while allowing for an outlier-free value model-based inference method.
Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models (2025.coling-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit impressive performance across various domains but struggle with arithmetic reasoning tasks.
Approach: They propose a Teaching-Inspired Integrated Prompting Framework which emulates the instructional process of a teacher guiding students.
Outcome: The proposed framework improves reasoning accuracy on nine benchmarks.
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning (2024.naacl-long)

Copied to clipboard

Challenge: TALMs have been successfully employed in question-answering benchmarks, but their efficacy on complex mathematical reasoning benchmarks are open research questions.
Approach: They propose a tool-augmented large language model for mathematical reasoning that enhances the skillset of large language models (LLMs) by 13.5%.
Outcome: The proposed model achieves better accuracy and better knowledge retrieval performance than existing tools.
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks (2024.acl-short)

Copied to clipboard

Challenge: Despite the generality and far-reaching consequences of large language models, there are still significant limitations making it difficult to apply them to certain tasks.
Approach: They show that large language models can perform arithmetic tasks more robustly when conditioned on all of the correct higher-order digits.
Outcome: The proposed model can predict the first digit of n-digit by m-digit multiplication without chain of thought reasoning, but in practice it fails to correctly predict the last digit on n digit by 1-digit multiplikation .
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction (2024.findings-acl)

Copied to clipboard

Challenge: Existing evaluations focus on problem-solving from examiner perspective, overlooking a dual perspective of examiner regarding error identification and correction.
Approach: They propose to use an annotated dataset to evaluate large language models from the examiner perspective and to use diverse prompts to evaluate eleven representative LLMs.
Outcome: The proposed model outperforms all models while LLaMA-2-7B has comparable abilities to closed-source models GPT-3.5 and Gemini Pro.
RankPrompt: Step-by-Step Comparisons Make Language Models Better Reasoners (2024.lrec-main)

Copied to clipboard

Challenge: Existing solutions to reasoning tasks require extensive human annotations or fail in scenarios with inconsistent responses.
Approach: They propose a new method that enables LLMs to self-rank their responses without additional resources.
Outcome: The proposed method improves reasoning performance of ChatGPT and GPT-4 with 13% improvement over existing methods.
ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to train large language models do not capture how humans learn to think.
Approach: They propose a method to fine-tune large language models for mathematical reasoning by using a text-infilling task that predicts masked equations from a given solution.
Outcome: Experiments on GSM8K, MATH, and GSM-Symbolic show that ClozeMath surpasses baseline Masked Thought in performance and robustness with two test-time scaling decoding algorithms, Beam Search and Chain-of-Thought decoding.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations