Steady diffusion and circulation matching fashions might characterize a robust various to autoregressive approaches for language modelling (LM), as they unlock a bunch of benefits at present reserved for steady modalities, together with accelerated sampling and tilting. Lately, a number of works have demonstrated the potential of producing discrete knowledge constantly by a easy circulation matching course of between a Gaussian and the one-hot encoded knowledge distribution. They’ve additional proven the feasibility of accelerated sampling by way of Categorical Move Maps (CFMs), leading to aggressive pattern high quality within the few-step regime. Nonetheless, this technique had solely been evaluated at comparatively modest scales (< 1B), leaving the query of its scalability fully open. On this article, we prepare a 1.7B-parameter base circulation mannequin on 2.1T tokens and self-distill it right into a CFM that generates numerous, high-quality textual content in as few as 4 inference steps whereas sustaining near-data-level token entropy. Moreover, we introduce a chance certain for CFMs within the semi-discrete setting, and present that they can be utilized to attain the mannequin on customary LM benchmarks, attaining ends in the identical vary as discrete diffusion strategies. Lastly, we uncover among the challenges that come up from coaching these fashions at scale, and we offer prescriptive insights on loss weighting and time scheduling.
† College of Oxford** Work achieved whereas at Apple

