Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond (2025.acl-industry)
Copied to clipboard
Liang Wen, Yunke Cai, Fenrui Xiao, Xin He, Qi An, Zhenyu Duan, Yimin Du, Junchen Liu, Tanglifu Tanglifu, Xiaowei Lv, Haosheng Zou, Yongchao Deng, Shousheng Jia, Xiangzheng Zhang
| Challenge: | Experimental results show that opensource curriculum training is more effective when distinct datasets are available for different training stages. |
| Approach: | They propose an opensource suite for training long reasoning models using publicdata and models. |
| Outcome: | The proposed model outperforms DeepSeek-R1-DistillQwen-32B models in math reasoning. |
Similar Papers
FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Improving training efficiency remains a challenge in large-scale Reinforcement Learning (RL). |
| Approach: | They propose a curriculum RL framework with stage-wise context scaling to improve RL training efficiency. |
| Outcome: | The proposed framework outperforms state-of-the-art reasoning models on five benchmarks and achieves 49.6% accuracy on AIME 2024. |
One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL (2025.acl-industry)
Copied to clipboard
Hyungjoo Chae, Dongjin Kang, Jihyuk Kim, Beong-woo Kwak, Sunghyun Park, Haeju Park, Jinyoung Yeo, Moontae Lee, Kyungjae Lee
| Challenge: | Existing large reasoning models are limited by their closed nature and high API costs and safety issues. |
| Approach: | They propose to build a long CoT dataset with existing short CoT LLMs that are not trained for inference-time scaling. |
| Outcome: | The proposed model achieves quality comparable to—or slightly below—R1 and is able to think longer and provide control over the thought budget to better manage the overthinking problem. |
Table-R1: Inference-Time Scaling for Table Reasoning Tasks (2025.emnlp-main)
Copied to clipboard
| Challenge: | In this study, we explore inference-time scaling on table reasoning tasks. |
| Approach: | They propose a large-scale dataset of reasoning traces and a reinforcement learning with verifiable rewards approach to enable inference-time scaling on table reasoning tasks. |
| Outcome: | The proposed model matches or exceeds GPT-4.1 and DeepSeek-R1 models on diverse table reasoning tasks. |
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models (2025.acl-long)
Copied to clipboard
Huifeng Yin, Yu Zhao, Minghao Wu, Xuanfan Ni, Bo Zeng, Huaiyu.wh Huaiyu.wh, Tianqi Shi, Liangying Shao, Chenyang Lyu, Longyue Wang, Weihua Luo, Kaifu Zhang
| Challenge: | Recent efforts to distill large reasoning models into smaller lightweight models have shown competitive performances. |
| Approach: | They propose to distill long Chain-of-Thought data to improve SFT and RL methods by constructing data from scratch using Monte Carlo Tree Search. |
| Outcome: | The proposed method significantly improves reasoning performance on various benchmarks such as math (GSM8K, MATH, AIME). |
Reflect, Rewrite, Repeat: How Simple Arithmetic Enables Advanced Reasoning in Small Language Models (2026.findings-eacl)
Copied to clipboard
Mengdie Flora Wang, Haochen Xie, Mun Young Kim, Baishali Chaudhury, Meghana Ashok, Suren Gunturu, Sungmin Hong, Jae Oh Woo
| Challenge: | Recent advances in language model reasoning require computationally intensive reinforcement learning and massive datasets. |
| Approach: | They propose a framework that combines Direct Preference Optimization and Supervised Fine-Tuning with selective guidance from larger models and iteratively refining solutions through a "reflect, rewrite, repeat" cycle. |
| Outcome: | The proposed framework shows significant performance improvements across arithmetic, symbolic and cognitive reasoning benchmarks. |
PARIF: Pushing the Pareto Frontier of Instruction Following and Reasoning with Curriculum Reinforcement Learning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing alignment methods struggle to balance general reasoning with instruction-following (IF) this is hindered by dependency on teacher models, reward hacking, and reasoning-answer inconsistencies. |
| Approach: | They propose a two-stage curriculum learning framework based on Reinforcement Learning from Verifiable Rewards to enhance both IF and general reasoning capabilities. |
| Outcome: | The proposed framework outperforms leading models on six representative IF tasks while achieving a 21.25% relative average improvement over the original model. |
Towards A Better Initial Policy Model For Scalable Long-CoT Reinforcement Learning (2025.findings-acl)
Copied to clipboard
| Challenge: | Long-CoT reasoning and reinforcement learning are demonstrating remarkable performance and scalability, however, there is a lack of systematic guidelines for obtaining a better initial policy model. |
| Approach: | They propose a systematic guideline and a novel Re-RFT method to obtain more efficient reasoning patterns from different initial models. |
| Outcome: | The proposed method surpasses DeepSeek-R1-Distill-Qwen-14B model by 4.6%, demonstrating its effectiveness and superiority. |
SLR: Automated Synthesis for Scalable Logical Reasoning (2026.acl-long)
Copied to clipboard
Lukas Helff, Ahmad Omar, Felix Friedrich, Antonia Wüst, Hikaru Shindo, Rupert Mitchell, Tim Woydt, Patrick Schramowski, Wolfgang Stammer, Kristian Kersting
| Challenge: | Existing benchmarks intended to evaluate reasoning capabilities emphasize deductive reasoning, where conclusions necessarily follow from given premises. |
| Approach: | They propose an end-to-end framework for systematic evaluation and training of Large Language Models via Scalable Logical Reasoning. |
| Outcome: | The proposed framework doubles Llama-3-8B accuracy on SLR-Bench, achieving parity with Gemini-Flash-Thinking at a fraction of computational cost. |
Predicate-Guided Generation for Mathematical Reasoning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that Prolog-MATH generates 81.3% solution coverage on Deepseek-V3 . |
| Approach: | They propose a curated corpus to support mathematical reasoning in large language models . they propose supervised fine-tuning followed by GRPO training to address problems that Deepseek-V3 fails to solve. |
| Outcome: | The proposed pipeline achieves 81.3% solution coverage on the Deepseek-V3 training set. |
When Internalization Fails: Finding Better Targets for Reasoning Compression (2026.findings-acl)
Copied to clipboard
| Challenge: | Reasoning language models generate long reasoning traces that increase latency and cost. |
| Approach: | They compare three approaches to shorten reasoning traces by inference-time truncation . they use Implicit Chain-of-Thought-style curricula that progressively shorten the teacher trace . |
| Outcome: | The proposed methods work well on GSM8K and multiplication tasks. |