Posts by Collection

notes

portfolio

publications

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) · arXiv preprint, 2026

A unified RLVR framework for dense image and video captioning, where caption quality is optimized through verifiable downstream question-answering rewards.

Recommended citation: Penghui Yang*, Long Xing*, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Yibin Wang, Yujie Zhou, Jiazi Bu, Jianze Liang, Qidong Huang, Jiaqi Wang, Feng Wu, and Dahua Lin. (2026). "CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning." arXiv preprint arXiv:2606.09393. Submitted to TPAMI.
Download Paper

ResOPD: Tail Residualization for Sparse On-Policy Distillation

arXiv Preprint · arXiv preprint, 2026

An exact coarse gradient plus a sampled fine-tail residual reduces gradient variance using only the teacher’s Top-k probabilities and the sampled-token score.

Recommended citation: Penghui Yang, Long Xing, Xuanlang Dai, Ziyu Liu, Kai Chen, and Yuhang Zang. (2026). "ResOPD: Tail Residualization for Sparse On-Policy Distillation." arXiv preprint arXiv:2610.04882.
Download Paper

talks

teaching

Teaching experience 1

Undergraduate course, University 1, Department, 2014

This is a description of a teaching experience. You can use markdown like any other post.

Teaching experience 2

Workshop, University 1, Department, 2015

This is a description of a teaching experience. You can use markdown like any other post.