View a PDF of the paper titled On the Place Bias of On-Coverage Distillation, by Yan Xie and 4 different authors
View PDF
HTML (experimental)
Summary:On-Coverage Distillation (OPD) improves the training effectivity of normal reinforcement studying by means of dense, token-level supervision from academics. In the usual KL goal of OPD, token-level losses are uniformly averaged, implying equal weights for all tokens. Nevertheless, we uncover that not all tokens are created equal: as pupil rollouts develop longer, they deviate farther from the trainer’s distribution, resulting in degraded supervision high quality at later positions. In consequence, OPD utilizing solely the primary 30% of tokens can carry out comparably to utilizing all tokens, whereas OPD utilizing solely the final 30% of tokens barely learns something. On this work, we offer a principled understanding of this challenge by means of the lens of constrained optimization. Primarily based on these insights, we derive Significance-Weighted On-Coverage Distillation (IW-OPD), during which the load assigned to every token is determined by the accrued discrepancy between the scholar’s and trainer’s distributions, naturally upweighting earlier tokens and downweighting later ones with bigger deviations. We present that IW-OPD converges considerably sooner than OPD, with higher studying effectivity, and achieves higher closing efficiency than normal OPD in each same-size and cross-scale settings, enhancing efficiency as much as 6.9 factors on AIME-2025.
Submission historical past
From: Yan Xie [view email] [v1]
Solar, 21 Jun 2026 17:20:21 UTC (202 KB)
[v2]
Tue, 23 Jun 2026 06:08:09 UTC (202 KB)
[v3]
Fri, 26 Jun 2026 02:45:32 UTC (202 KB)
![[2606.22600] On the Place Bias of On-Coverage Distillation [2606.22600] On the Place Bias of On-Coverage Distillation](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)
