ResOPD: Tail Residualization for Sparse On-Policy Distillation

Published in arXiv preprint, 2026

Penghui Yang, Long Xing, Xuanlang Dai, Ziyu Liu, Kai Chen, and Yuhang Zang

Paper | PDF | Code | Dataset

ResOPD overview: sparse teacher feedback, gradient variance reduction, and reasoning performance

ResOPD uses the teacher’s Top-k probabilities and the score of the token sampled by the student to estimate the full-vocabulary reverse-KL gradient. It computes the observable coarse gradient exactly and samples only the unresolved within-tail residual, without additional teacher forward passes.

Across five frozen-prefix panels, ResOPD reduces covariance trace by 52.5–75.0% relative to sampled-token on-policy distillation. For mathematical reasoning with Qwen3.5-27B distilled into Qwen3.5-4B, it achieves an average score of 50.12%, compared with 45.72% for HETS under the paper’s protocol.

The implementation builds on verl. The released DeepMath subset contains 1,024 training problems and 252 held-out problems in verl-compatible Parquet format.

Recommended citation: Penghui Yang, Long Xing, Xuanlang Dai, Ziyu Liu, Kai Chen, and Yuhang Zang. (2026). "ResOPD: Tail Residualization for Sparse On-Policy Distillation." arXiv preprint arXiv:2610.04882.
Download Paper