Challenge: Large language models (LLMs) can solve problems step-by-step, but it is unclear whether they know when to use CoT and whether they are always necessary.
Approach: They propose to use LLMs to generate redundant calculations and reasoning on a manually constructed math QA dataset, GSM8K-Zero.
Outcome: The proposed model generates redundant calculations and reasoning on a manually constructed math QA dataset, but it is unclear whether it is necessary to use CoT reasoning.

Similar Papers

LLMs Faithfully and Iteratively Compute Answers During CoT: A Systematic Analysis With Multi-step Arithmetics (2026.findings-eacl)

Copied to clipboard

Challenge: Specifically, we examine when the LLMs’ answer is (pre)determined, especially before the CoT begins or after, and how strongly the information from CoT specifically has a causal effect on the final answer.
Approach: They examine when the LLMs’ answer is (pre)determined, especially before the CoT begins or after, and how strongly the information from CoT specifically has a causal effect on the final answer.
Outcome: The proposed model can generate reasoning chains while generating the reasoning chain on the fly.
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning? (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in o1-like models have generated long Chain-of-Thought reasoning steps to improve the reasoning abilities of existing Large Language Models (LLMs).
Approach: They propose a DeltaBench to analyze the quality and effectiveness of o1-like models and measure their ability to detect errors in long COT reasoning.
Outcome: The proposed model can detect errors in long COT reasoning.
How Likely Do LLMs with CoT Mimic Human Reasoning? (2025.coling-main)

Copied to clipboard

Challenge: Using chain-of-thought to elicit reasoning capabilities is not always effective and accurate.
Approach: They compare the reasoning process of LLMs with humans to understand the causal chain . they find that LLM deviates from the ideal causal chain, resulting in spurious correlations .
Outcome: The proposed method does not improve performance or accurately represent reasoning processes in LLMs.
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are often evaluated on math word problems . however, such metrics conflate two distinct sub-skills: abstract formulation and arithmetic computation.
Approach: They propose to use Final-answer-based metrics to evaluate large language models on math word problems to conflate two distinct sub-skills: abstract formulation and arithmetic computation.
Outcome: The proposed model performance is bottlenecked by arithmetic computation and not abstract formulation, the study shows.
One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL (2025.acl-industry)

Copied to clipboard

Challenge: Existing large reasoning models are limited by their closed nature and high API costs and safety issues.
Approach: They propose to build a long CoT dataset with existing short CoT LLMs that are not trained for inference-time scaling.
Outcome: The proposed model achieves quality comparable to—or slightly below—R1 and is able to think longer and provide control over the thought budget to better manage the overthinking problem.
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) achieve impressive performance on complex benchmarks yet sometimes fail on basic math reasoning.
Approach: They propose a benchmark to evaluate the efficiency of reasoning in large language models . they formalize the accuracy-verbosity tradeoff and introduce the overthinking score .
Outcome: The proposed model performs well on complex benchmarks but fails on basic math reasoning . the proposed model generates 18 more tokens while achieving lower accuracy .
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction (2024.findings-acl)

Copied to clipboard

Challenge: Existing evaluations focus on problem-solving from examiner perspective, overlooking a dual perspective of examiner regarding error identification and correction.
Approach: They propose to use an annotated dataset to evaluate large language models from the examiner perspective and to use diverse prompts to evaluate eleven representative LLMs.
Outcome: The proposed model outperforms all models while LLaMA-2-7B has comparable abilities to closed-source models GPT-3.5 and Gemini Pro.
Memory or Reasoning? Explore How LLMs Compute Mixed Arithmetic Expressions (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) can solve complex multi-step math reasoning problems, but their internal implementation is limited.
Approach: They propose to use a "C**ausal **E**ffect **D**riven **F**ine-tuning method" to improve LLMs' reasoning ability.
Outcome: The proposed method improves the model's reasoning ability by enhancing key components that are used to execute mixed arithmetic calculations.
Working Memory Identifies Reasoning Limits in Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, we examine the limitations of their cognitive capabilities and their working memory.
Approach: They examine the limitations of large language models from a scaling perspective . they also assess various prompting strategies, revealing their diverse impacts on LLM performance.
Outcome: The proposed models perform poorly on n-back tasks and on prompting strategies.
Evaluating the Deductive Competence of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing large language models have limited abilities to solve deductive reasoning problems . performance differences between conditions do not improve overall performance .
Approach: They investigate whether several large language models can solve a deductive reasoning problem in their conventional form.
Outcome: The proposed models can solve a classic type of deductive reasoning problem in their conventional form.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations