{"id":4085,"date":"2026-08-20T00:00:00","date_gmt":"2026-08-20T00:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/"},"modified":"2026-08-22T09:59:04","modified_gmt":"2026-08-22T09:59:04","slug":"scaling-laws-mixture-pretraining","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/","title":{"rendered":"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p>As language fashions scale, the quantity of information they require grows \u2013 but many goal knowledge sources, similar to low-resource languages or specialised domains, are inherently restricted in dimension. A typical technique is to combine this scarce however worthwhile goal knowledge with plentiful generic knowledge, which presents a elementary trade-off: too little goal knowledge within the combination underexposes the mannequin to the goal area, whereas an excessive amount of goal knowledge repeats the identical examples excessively, yielding diminishing returns and eventual overfitting. We examine this trade-off throughout greater than 2,000 language-model coaching runs spanning a number of mannequin and goal dataset sizes, in addition to a number of knowledge varieties, together with multilingual, domain-specific, and quality-filtered mixtures. Throughout all settings, we discover that repetition is a central driver of target-domain efficiency, and that combination coaching tolerates a lot larger repetition than single-source coaching: scarce goal corpora could be reused 15\u201320 occasions, with the optimum variety of repetitions relying on the goal knowledge dimension, compute finances, and mannequin scale. Subsequent, we introduce a repetition-aware combination scaling legislation that accounts for the lowering worth of repeated goal tokens and the regularizing function of generic knowledge. Optimizing the scaling legislation supplies a principled technique to compute efficient combination configurations, yielding sensible combination suggestions for pretraining underneath knowledge constraints.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/machinelearning.apple.com\/research\/scaling-laws-mixture-pretraining\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>As language fashions scale, the quantity of information they require grows \u2013 but many goal knowledge sources, similar to low-resource languages or specialised domains, are inherently restricted in dimension. A typical technique is to combine this scarce however worthwhile goal knowledge with plentiful generic knowledge, which presents a elementary trade-off: too little goal knowledge within [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":4087,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[2],"tags":[1804,160,2615,4348,3114,325],"class_list":["post-4085","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-research-breakthroughs","tag-constraints","tag-data","tag-laws","tag-mixture","tag-pretraining","tag-scaling"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints - Future News 24<\/title>\n<meta name=\"description\" content=\"As language models scale, the amount of data they require grows \u2013 yet many target data sources, such as low-resource languages or\u2026\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints - Future News 24\" \/>\n<meta property=\"og:description\" content=\"As language models scale, the amount of data they require grows \u2013 yet many target data sources, such as low-resource languages or\u2026\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-20T00:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-22T09:59:04+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"1 minute\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints\",\"datePublished\":\"2026-08-20T00:00:00+00:00\",\"dateModified\":\"2026-08-22T09:59:04+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/\"},\"wordCount\":225,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"keywords\":[\"constraints\",\"data\",\"laws\",\"Mixture\",\"Pretraining\",\"Scaling\"],\"articleSection\":[\"AI Research &amp; Breakthroughs\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/\",\"name\":\"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"datePublished\":\"2026-08-20T00:00:00+00:00\",\"dateModified\":\"2026-08-22T09:59:04+00:00\",\"description\":\"As language models scale, the amount of data they require grows \u2013 yet many target data sources, such as low-resource languages or\u2026\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"contentUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/20\\\/scaling-laws-mixture-pretraining\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints - Future News 24","description":"As language models scale, the amount of data they require grows \u2013 yet many target data sources, such as low-resource languages or\u2026","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/","og_locale":"en_US","og_type":"article","og_title":"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints - Future News 24","og_description":"As language models scale, the amount of data they require grows \u2013 yet many target data sources, such as low-resource languages or\u2026","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/","og_site_name":"Future News 24","article_published_time":"2026-08-20T00:00:00+00:00","article_modified_time":"2026-08-22T09:59:04+00:00","og_image":[{"url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"1 minute"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints","datePublished":"2026-08-20T00:00:00+00:00","dateModified":"2026-08-22T09:59:04+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/"},"wordCount":225,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","keywords":["constraints","data","laws","Mixture","Pretraining","Scaling"],"articleSection":["AI Research &amp; Breakthroughs"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/","name":"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","datePublished":"2026-08-20T00:00:00+00:00","dateModified":"2026-08-22T09:59:04+00:00","description":"As language models scale, the amount of data they require grows \u2013 yet many target data sources, such as low-resource languages or\u2026","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/#primaryimage","url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","contentUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/20\/scaling-laws-mixture-pretraining\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Scaling Legal guidelines for Combination Pretraining Underneath Information Constraints"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4085","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=4085"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4085\/revisions"}],"predecessor-version":[{"id":4086,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4085\/revisions\/4086"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/4087"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=4085"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=4085"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=4085"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}