Challenge: Program-of-Thought (PoT) replaces natural language-based Chain-ofThough (CoT) but introduces more reasoning errors, such as incorrect formulas or flawed logic, compared to CoT.
Approach: They propose a method that integrates CoT and Program-of-Thought to achieve more accurate reasoning and reinforcement learning.
Outcome: The proposed method achieves an average improvement of 6.5% on the Llama-Base model and 4.3% on the Mistral-Bass model across 8 mathematical calculation datasets.

Similar Papers

Chain-of-Thought Reasoning in Tabular Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to extend chain-of-thought reasoning into large language models are not viable in the scenario of privatization deployment or limited resources.
Approach: They propose a framework that extends chain-of-thought reasoning into tabular language models . framework coordinates two TaLMs responsible for CoT generation and answer inference .
Outcome: The proposed framework outperforms the state-of-the-art ChatGPT on the TABMWP dataset by 9.55% (82.60%92.15% in accuracy) with less parameters (0.8B).
Think Like You Execute: Verifiable Chain of Thought from Program Traces (2026.acl-industry)

Copied to clipboard

Challenge: Current synthetic Chain-of-Thought (CoT) training data often consists of plausible-sounding explanations generated by teacher models, not verifiable accounts of actual program behavior.
Approach: They propose to ground CoT generation directly in program execution traces to improve reasoning capabilities.
Outcome: The proposed pipeline improves performance on live code benchmarks and on cruxEval-output and cruxeval-input.
Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work focuses on enabling models to generate natural language chain-of-thought rationales or leverage executable and verifiable code, such as Python.
Approach: They propose a novel training pipeline that integrates sequential P-CoT and N-Co T generation and a subtask hybrid training strategy to facilitate natural language transferability.
Outcome: The proposed training pipeline improves both N-CoT and P-Co T performance over the RL baseline.
How Likely Do LLMs with CoT Mimic Human Reasoning? (2025.coling-main)

Copied to clipboard

Challenge: Using chain-of-thought to elicit reasoning capabilities is not always effective and accurate.
Approach: They compare the reasoning process of LLMs with humans to understand the causal chain . they find that LLM deviates from the ideal causal chain, resulting in spurious correlations .
Outcome: The proposed method does not improve performance or accurately represent reasoning processes in LLMs.
Wrong-of-Thought: An Integrated Reasoning Framework with Multi-Perspective Verification and Wrong Information (2024.findings-emnlp)

Copied to clipboard

Challenge: Chain-of-Thought (CoT) is a key technique for enhancing the performance of Large Language Models.
Approach: They propose a framework that optimizes outputs by utilizing wrong information and multi-perspective verification.
Outcome: The proposed framework surpasses all baselines on 8 datasets and 5 LLMs.
Markov Chain of Thought for Efficient Mathematical Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies have sought to enhance the mathematical reasoning capabilities of large language models.
Approach: They propose a Markov Chain of Thought (MCoT) that compresses previous reasoning steps into a simplified question.
Outcome: The proposed method improves efficiency and maintains comparable accuracy.
GoT: Effective Graph-of-Thought Reasoning in Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have been advancing at an unprecedented pace.
Approach: They propose a graph-based approach which models human thought processes as a chain and as 'graphs' by representing thought units as nodes and connections between them as edges, they capture the non-sequential nature of human thinking and allows for a more realistic modeling of thought processes.
Outcome: The proposed model improves on a text-only reasoning task and a multimodal reasoning task.
Program-of-Thought Reveals LLM Abstraction Ceilings (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models exhibit reasoning ability when supervised with chain-of-thought (CoT) traces.
Approach: They evaluate large language models with CoT traces and fine-tune them with Program-of-Thought supervision.
Outcome: The proposed model performance degrades sharply under numeric perturbations under isomorphic variants.
CoT-RAG: Integrating Chain of Thought and Retrieval-Augmented Generation to Enhance Reasoning in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Chain-of-thought reasoning has two key limitations: lack of reliability when solely relying on LLM-generated reasoning chains and interference from natural language reasoning steps with the models’ inference logic.
Approach: They propose a chain-of-thought reasoning framework with three key designs to address these issues.
Outcome: The proposed framework improves the performance of large language models on complex tasks by incorporating knowledge graphs and learnable knowledge case-aware RAG.
Over-Reasoning and Redundant Calculation of Large Language Models (2024.eacl-short)

Copied to clipboard

Challenge: Large language models (LLMs) can solve problems step-by-step, but it is unclear whether they know when to use CoT and whether they are always necessary.
Approach: They propose to use LLMs to generate redundant calculations and reasoning on a manually constructed math QA dataset, GSM8K-Zero.
Outcome: The proposed model generates redundant calculations and reasoning on a manually constructed math QA dataset, but it is unclear whether it is necessary to use CoT reasoning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations