On-policy distillation gives dense, per-token supervision for coaching reasoning fashions; nevertheless, it stays unclear underneath which situations this sign is useful and underneath which it’s detrimental. Which trainer mannequin ought to be used, and within the case of self-distillation, which particular context ought to function the supervisory sign? Does the optimum alternative range from one token to the subsequent? At current, addressing these questions usually requires expensive coaching runs whose combination efficiency metrics obscure the dynamics on the degree of particular person tokens. We introduce a training-free diagnostic framework that operates on the highest decision: per token, per query, and per trainer. We derive an excellent per-node gradient outlined because the parameter replace that maximally will increase the coed’s chance of success. We then develop a scalable targeted-rollout algorithm to estimate this gradient effectively, even for lengthy chains of intermediate ideas. The gradient alignment rating, outlined because the cosine similarity between this excellent gradient and any given distillation gradient, quantifies the extent to which a specific configuration approximates the perfect sign. Throughout a variety of self-distillation settings and exterior trainer fashions, we observe that distillation steerage displays considerably greater alignment with the perfect on incorrect rollouts than on right ones, the place the coed already performs effectively and the trainer’s sign tends to turn out to be noisy. Moreover, we discover that the optimum distillation context relies upon collectively on the coed mannequin’s capability and the goal job, and that no single universally efficient configuration emerges. These findings inspire the usage of per-task, per-token diagnostic analyses for distillation.

