{"id":2630,"date":"2026-07-20T00:00:00","date_gmt":"2026-07-20T00:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/"},"modified":"2026-07-21T02:59:21","modified_gmt":"2026-07-21T02:59:21","slug":"length-value-model","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/","title":{"rendered":"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p>Token serves as the basic unit of computation in trendy autoregressive fashions, and era size straight influences each inference price and reasoning efficiency. Regardless of its significance, present approaches lack fine-grained size modeling, working primarily on the coarse-grained sequence degree. On this paper, we introduce the Size Worth Mannequin (LenVM), a token-level framework that fashions the remaining era size at every decoding step. By formulating size modeling as a worth estimation drawback and assigning a relentless adverse reward to every generated token, LenVM predicts a bounded, discounted return that serves as a monotone proxy for the remaining era horizon. This formulation yields supervision that&#8217;s annotation-free, dense, unbiased, and scalable. Experiments on LLMs and VLMs exhibit that LenVM gives a extremely efficient sign at inference time. On the LIFEBench precise size matching process, making use of LenVM to a 7B mannequin improves the size rating from 30.9 to 64.8, considerably outperforming frontier closed-source fashions. Moreover, LenVM allows steady management over the commerce off between efficiency and effectivity. On GSM8K at a finances of 200 tokens, LenVM maintains 63 % accuracy in comparison with 6 % for token finances baseline. It additionally precisely predicts complete era size from the immediate boundary. Lastly, LenVM\u2019s token-level values provide an interpretable view of era dynamics, revealing how particular tokens shift reasoning towards shorter or longer regimes. Outcomes exhibit that LenVM helps a broad vary of purposes, together with size management, prediction, and interpretation of era dynamics. They counsel that era size might be successfully modeled as a token-level worth sign, highlighting the potential of LenVM as a normal framework for size modeling and as a length-specific worth sign that would assist future RL coaching.<\/p>\n<p>\u2020 College of California, Santa Barbara\u2021 Carnegie Mellon College\u00a7 LMSYS Org\u00b6 College of Wisconsin\u2013Madison<\/p><\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/machinelearning.apple.com\/research\/length-value-model\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Token serves as the basic unit of computation in trendy autoregressive fashions, and era size straight influences each inference price and reasoning efficiency. Regardless of its significance, present approaches lack fine-grained size modeling, working primarily on the coarse-grained sequence degree. On this paper, we introduce the Size Worth Mannequin (LenVM), a token-level framework that fashions [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":2632,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[2],"tags":[3112,105,1938,3114,3113,3115],"class_list":["post-2630","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-research-breakthroughs","tag-length","tag-model","tag-modeling","tag-pretraining","tag-scalable","tag-tokenlevel"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling - Future News 24<\/title>\n<meta name=\"description\" content=\"Token serves as the fundamental unit of computation in modern autoregressive models, and generation length directly influences both\u2026\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Token serves as the fundamental unit of computation in modern autoregressive models, and generation length directly influences both\u2026\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-20T00:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-21T02:59:21+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"1 minute\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling\",\"datePublished\":\"2026-07-20T00:00:00+00:00\",\"dateModified\":\"2026-07-21T02:59:21+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/\"},\"wordCount\":303,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"keywords\":[\"Length\",\"Model\",\"Modeling\",\"Pretraining\",\"Scalable\",\"TokenLevel\"],\"articleSection\":[\"AI Research &amp; Breakthroughs\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/\",\"name\":\"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"datePublished\":\"2026-07-20T00:00:00+00:00\",\"dateModified\":\"2026-07-21T02:59:21+00:00\",\"description\":\"Token serves as the fundamental unit of computation in modern autoregressive models, and generation length directly influences both\u2026\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"contentUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/20\\\/length-value-model\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling - Future News 24","description":"Token serves as the fundamental unit of computation in modern autoregressive models, and generation length directly influences both\u2026","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/","og_locale":"en_US","og_type":"article","og_title":"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling - Future News 24","og_description":"Token serves as the fundamental unit of computation in modern autoregressive models, and generation length directly influences both\u2026","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/","og_site_name":"Future News 24","article_published_time":"2026-07-20T00:00:00+00:00","article_modified_time":"2026-07-21T02:59:21+00:00","og_image":[{"url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"1 minute"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling","datePublished":"2026-07-20T00:00:00+00:00","dateModified":"2026-07-21T02:59:21+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/"},"wordCount":303,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","keywords":["Length","Model","Modeling","Pretraining","Scalable","TokenLevel"],"articleSection":["AI Research &amp; Breakthroughs"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/","name":"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","datePublished":"2026-07-20T00:00:00+00:00","dateModified":"2026-07-21T02:59:21+00:00","description":"Token serves as the fundamental unit of computation in modern autoregressive models, and generation length directly influences both\u2026","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/#primaryimage","url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","contentUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/20\/length-value-model\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Size Worth Mannequin: Scalable Worth Pretraining for Token-Degree Size Modeling"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2630","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=2630"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2630\/revisions"}],"predecessor-version":[{"id":2631,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2630\/revisions\/2631"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/2632"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=2630"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=2630"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=2630"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}