Can a big language mannequin (LLM) enhance at code era utilizing solely its personal uncooked outputs, and not using a verifier, a trainer mannequin, or reinforcement studying? We reply within the affirmative with easy self-distillation (SSD): pattern options from the mannequin with sure temperature and truncation configurations, then fine-tune on these samples with normal supervised fine-tuning. SSD improves Qwen3-30B-Instruct from 42.4% to 55.3% move@1 on LiveCodeBench v6, with features concentrating on tougher issues, and it generalizes throughout Qwen and Llama fashions at 4B, 8B, and 30B scale, together with each instruct and considering variants. To grasp why such a easy technique can work, we hint these features to a precision-exploration battle in LLM decoding and present that SSD reshapes token distributions in a context-dependent approach, suppressing distractor tails the place precision issues whereas preserving helpful range the place exploration issues. Taken collectively, SSD gives a complementary post-training route for bettering LLM code era.

![[2606.11056] On pseudogap section as precursor to a superconducting dome in high-Tc cuprates: Non-analytic T* as a perform of doping [2606.11056] On pseudogap section as precursor to a superconducting dome in high-Tc cuprates: Non-analytic T* as a perform of doping](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)