On-Policy Distillation Teaches New Skills but Not New Knowledge
Abstract
On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning. Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion. Together, these results demonstrate that on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.
Community
TL;DR: On-policy distillation (OPD) does not expand parametric memory—it teaches models how to organize what they already know.
Key Findings:
- Skill vs. Knowledge: Reverse-KL OPD reliably transfers compositional reasoning across unseen structures, but transfers minimal factual knowledge.
- The Mechanism: Switching from reverse-KL to forward-KL restores factual transfer, whereas student rollouts specifically drive multi-step reasoning execution.
- Empirical Validation: Tested across 4 model families: yields strong gains on competition math, but zero memory expansion on factual QA.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning (2026)
- Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives (2026)
- DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation (2026)
- CA-OPD: Confidence-Aware On-Policy Distillation for Structured Visual Prediction (2026)
- SAPD: Step-Aligned Privileged Distillation (2026)
- Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation (2026)
- Recursive Self-Improvement via On-Policy Distillation for Reasoning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper