Challenge: Existing approaches to incentivize LLMs’ deep thinking abilities require large-scale data or significant training efforts.
Approach: They introduce an efficient framework that enhances LLM reasoning by teaching models to self-verify and self-correct during inference.
Outcome: The proposed framework outperforms models trained on long-CoT distilled data with 3.1k initialization samples and achieves an accuracy improvement of 51.0% to 81.6%.

Similar Papers

Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models have demonstrated notable inferential capacities via reinforcement learning (RL) however, “zero-RL” approaches relying on fixed prompt templates introduce substantial sampling inefficiencies for weak LLMs.
Approach: They propose a hierarchical metacognitive RL framework that decomposes zero-accuracy problems into subproblems and prompts the policy to refine answers by referencing previous wrong solutions.
Outcome: The proposed framework improves sample utilization and sample efficiency and accelerates convergence compared to baselines.
Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable abilities, but they invariably generate flawed responses.
Approach: They propose a self-correction approach that instructs VLMs to refine their outputs by allowing them to learn from their self-generated self-reference data without external feedback.
Outcome: The proposed approach enables VLMs to learn from their self-generated self-correction data without relying on external feedback, facilitating self-improvement.
Small Language Models Need Strong Verifiers to Self-Correct Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies show that large language models can self-correct their outputs by generating a critique and revising it based on the critique.
Approach: They propose a pipeline that prompts small language models to collect self-correction data that supports the training of self-refinement abilities.
Outcome: The proposed pipeline improves the self-correction abilities of two models on five datasets spanning math and commonsense reasoning.
Large Language Models Can Self-Improve (2023.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have excellent performance in various tasks, but fine-tuning requires extensive supervision.
Approach: They propose to use a pre-trained Large Language Model to generate rationale-augmented answers for unlabeled questions and fine-tune the LLM using those self-generated solutions as target outputs.
Outcome: The proposed approach improves the general reasoning ability of a 540B-parameter LLM without any ground truth label.
ReActR: Reasoning through Error-Activated Reflection for LLM Post-Training (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for improving the mathematical abilities of Large Language Models (LLMs) focus disproportionately on scaling correct training samples, overlooking the rich learning signals contained in erroneous reasoning trajectories.
Approach: They propose a framework that enhances reasoning by learning reflective behaviors from erroneous trajectories by using data construction and training.
Outcome: Extensive experiments on three LLMs show that ReActR improves reasoning performance on Llama-3-8B.
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains (2026.acl-long)

Copied to clipboard

Challenge: Recent large language models (LLMs) have demonstrated remarkable progress in reasoning, but their applications on knowledge-intensive domains have not been explored due to the scarcity of high-quality verifiable data.
Approach: They propose a framework that extends reinforcement learning with verifiable rewards (RLVR) to knowledge-intensive domains through automated verififiability data synthesis while enabling verification of the LLM's reasoning process.
Outcome: Extensive experiments show that the proposed framework enhances the reasoning of large language models in knowledge-intensive domains without significantly compromising the model’s general capabilities.
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on correctness, overlooking optimality . large language models excel at math, coding, logic and puzzles .
Approach: They propose a framework for training and evaluating Large Language Models on NP-hard optimization problems through quality-aware RLVR.
Outcome: The proposed framework outperforms existing benchmarks on math, coding, logic and puzzles.
Correct, Concise and Complete: Multi-stage Training For Adaptive Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) increase test-time computation, often in the form of chain-of-thought (CoT) however, reasoning traces can become unnecessarily long, increasing computation costs without improving accuracy and sometimes even degrading performance.
Approach: They propose a multi-stage efficient reasoning method that combines supervised fine-tuning with reinforcement learning using an adaptive length penalty.
Outcome: The proposed method reduces response length by an average of 28% for 8B models and 40% for 32B models while incurring only minor performance drops of 1.6 and 2.5 points, respectively.
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement (2024.findings-acl)

Copied to clipboard

Challenge: Pretrained large language models are currently state-of-the-art for solving most tasks . however, many of them are in the low-data regime, making fine-tuning challenging . a new data augmentation strategy uses a teacher LLM to augment a small seed dataset .
Approach: They propose a targeted and iterative data augmentation strategy that augments a teacher LLM to fine-tune a small seed dataset by adding additional data.
Outcome: The proposed approach outperforms fine-tuning and other data augmentation strategies on a small seed dataset.
Steering LLM Reasoning Through Bias-Only Adaptation (2025.emnlp-main)

Copied to clipboard

Challenge: Compared with LoRA and BitFit, training a single steering vector per layer with reinforcement learning requires orders of magnitude fewer resources and isolates a much smaller, more interpretable parameter set.
Approach: They propose to train a single steering vector per layer with reinforcement learning while freezing all base weights to match the accuracy of fully RL-tuned reasoning models.
Outcome: The proposed approach improves on an 8 billion-parameter model while keeping all base weights fixed.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations