As language fashions scale, the quantity of information they require grows – but many goal knowledge sources, similar to low-resource languages or specialised domains, are inherently restricted in dimension. A typical technique is to combine this scarce however worthwhile goal knowledge with plentiful generic knowledge, which presents a elementary trade-off: too little goal knowledge within the combination underexposes the mannequin to the goal area, whereas an excessive amount of goal knowledge repeats the identical examples excessively, yielding diminishing returns and eventual overfitting. We examine this trade-off throughout greater than 2,000 language-model coaching runs spanning a number of mannequin and goal dataset sizes, in addition to a number of knowledge varieties, together with multilingual, domain-specific, and quality-filtered mixtures. Throughout all settings, we discover that repetition is a central driver of target-domain efficiency, and that combination coaching tolerates a lot larger repetition than single-source coaching: scarce goal corpora could be reused 15–20 occasions, with the optimum variety of repetitions relying on the goal knowledge dimension, compute finances, and mannequin scale. Subsequent, we introduce a repetition-aware combination scaling legislation that accounts for the lowering worth of repeated goal tokens and the regularizing function of generic knowledge. Optimizing the scaling legislation supplies a principled technique to compute efficient combination configurations, yielding sensible combination suggestions for pretraining underneath knowledge constraints.

![[2606.07559] Phantom Transitions in Language Mannequin Superb-Tuning: A Density-Matrix Evaluation [2606.07559] Phantom Transitions in Language Mannequin Superb-Tuning: A Density-Matrix Evaluation](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)