Challenge: a step-by-step reasoning approach like chain of thought has proved to be effective in eliciting reasoning abilities in large language models.
Approach: They propose a knowledge distillation approach that leverages CoT reasoning capabilities of larger models and distills them into smaller models.
Outcome: The proposed scheme boosts the performance of smaller models over 70% on multiple reasoning datasets.

Similar Papers

PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning (2024.naacl-long)

Copied to clipboard

Challenge: Large language models excel in various tasks, but their huge size and inaccessibility of parameters present challenges for practical deployment.
Approach: They propose to use CoT data to distill task-specific ability from large language models to smaller models . they use reasoning programs to suppress errors in distilled data and improve distillation quality .
Outcome: The proposed model outperforms LLMs on arithmetic reasoning, symbolic reasoning, and general ability.
Mixed Distillation Helps Smaller Language Models Reason Better (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent large language models (LLMs) have demonstrated impressive multiple step-by-step reasoning capabilities in recent NLP reasoning tasks.
Approach: They propose a mixed distillation framework that distills multiple step-by-step reasoning abilities into smaller language models (SLMs) they leverage LLMs to generate multiple step by step reasoning rationales by sampling automatically.
Outcome: The proposed framework outperforms existing models on SVAMP, GSM8K and ASDIV, while a single model generated by MD exceeds the comprehensive performance of two individual CoT and PoT distilled models.
Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key Information (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods focus on transferring teacher-generated rationales to student models, but do not explore teachers’ dynamic attention towards critical information during reasoning.
Approach: They propose a method that transfers the teacher’s stepwise attention on key information to the student model and a Mixture of Layers module that allows dynamic alignment between the teacher and student.
Outcome: The proposed framework achieves consistent performance improvements across multiple mathematical and commonsense reasoning datasets.
MoDE-CoTD: Chain-of-Thought Distillation for Complex Reasoning Tasks with Mixture of Decoupled LoRA-Experts (2024.lrec-main)

Copied to clipboard

Challenge: Current Chain-of-thought Distillation methods hinder CoT reasoning performance . student models are separately distilled from specific reasoning tasks . parameter update of student models severely harms CoT ability on unseen reasoning tasks.
Approach: They propose a method which distills Chain-of-thought reasoning ability of large language models to much smaller student models.
Outcome: The proposed method improves the reasoning ability of large language models on 14 datasets.
Small Models Struggle to Learn from Strong Reasoners (2025.findings-acl)

Copied to clipboard

Challenge: a small learning gap exists between large and small language models . long CoT data and large model responses are not beneficial for small models - a problem that may be due to the small student model's ability to handle distribution shifts.
Approach: They propose a mix distillation strategy that balances reasoning complexity by combining long and short CoT examples or reasoning from both larger and smaller models.
Outcome: The proposed strategy outperforms training on large and small models on short CoT and small model CoT.
Probe Then Retrieve and Reason: Distilling Probing and Reasoning Capabilities into Smaller Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent research efforts have focused on distilling Large Language Models into Small Language Model (SLMs) however, the results of CoT distillation are inadequate for knowledge-intensive reasoning tasks.
Approach: They propose a retrieval-based framework which distills question probing and reasoning capabilities from Large Language Models into SLMs.
Outcome: The proposed framework improves probing and reasoning capabilities of large language models in knowledge-intensive reasoning tasks.
Capture the Key in Reasoning to Enhance CoT Distillation Generalization (2025.acl-long)

Copied to clipboard

Challenge: Existing distillation methods for Large Language Models (LLMs) focus on fine-tuning student SLMs on correct data, resulting in students struggling to learn the key instead of analyzing mistakes according to correct solutions.
Approach: They propose a method that exposes key reasoning steps rather than simple fine-tuning students' CoTs data by using a set of prompts with similar reasoning paths but divergent conclusions.
Outcome: The proposed method improves student SLMs' ability to learn key reasoning steps rather than fine-tuning them on teacher data.
Mentor-KD: Making Small Language Models Better Multi-step Reasoners (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown impressive emergent capabilities by leveraging Chain-of-Thought (CoT) prompting.
Approach: They propose a Knowledge Distillation approach which transfers multi-step reasoning ability of Large Language Models (LLMs) to smaller LMs by fine-tuning language models of multi- step rationales generated by LLM teachers.
Outcome: The proposed method is able to transfer multi-step reasoning ability of LLMs to smaller LMs while addressing data quality and soft label provision.
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models (2025.acl-long)

Copied to clipboard

Challenge: Recent efforts to distill large reasoning models into smaller lightweight models have shown competitive performances.
Approach: They propose to distill long Chain-of-Thought data to improve SFT and RL methods by constructing data from scratch using Monte Carlo Tree Search.
Outcome: The proposed method significantly improves reasoning performance on various benchmarks such as math (GSM8K, MATH, AIME).
Learning to Maximize Mutual Information for Chain-of-Thought Distillation (2024.findings-acl)

Copied to clipboard

Challenge: Knowledge distillation is a technique of transferring knowledge from large, complex models to smaller ones.
Approach: They propose a method utilizing chain-of-thought distillation to transfer knowledge from large, complex models to smaller ones by maximizing mutual information of the representation features of the two tasks.
Outcome: The proposed method outperforms the state-of-the-art knowledge distillation method on four datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations