ResOPD: Tail Residualization for Sparse On-Policy Distillation
An exact coarse gradient plus a sampled fine-tail residual reduces gradient variance using only the teacher’s Top-k probabilities and the sampled-token score.
An exact coarse gradient plus a sampled fine-tail residual reduces gradient variance using only the teacher’s Top-k probabilities and the sampled-token score.
A unified RLVR framework for dense image and video captioning, where caption quality is optimized through verifiable downstream question-answering rewards.