ResOPD: Tail Residualization for Sparse On-Policy Distillation
Published in arXiv preprint, 2026
Penghui Yang, Long Xing, Xuanlang Dai, Ziyu Liu, Kai Chen, and Yuhang Zang

ResOPD uses the teacher’s Top-k probabilities and the score of the token sampled by the student to estimate the full-vocabulary reverse-KL gradient. It computes the observable coarse gradient exactly and samples only the unresolved within-tail residual, without additional teacher forward passes.
Across five frozen-prefix panels, ResOPD reduces covariance trace by 52.5–75.0% relative to sampled-token on-policy distillation. For mathematical reasoning with Qwen3.5-27B distilled into Qwen3.5-4B, it achieves an average score of 50.12%, compared with 45.72% for HETS under the paper’s protocol.
The implementation builds on verl. The released DeepMath subset contains 1,024 training problems and 252 held-out problems in verl-compatible Parquet format.
Recommended citation: Penghui Yang, Long Xing, Xuanlang Dai, Ziyu Liu, Kai Chen, and Yuhang Zang. (2026). "ResOPD: Tail Residualization for Sparse On-Policy Distillation." arXiv preprint arXiv:2610.04882.
Download Paper
