{"id":3560,"date":"2026-08-07T00:00:00","date_gmt":"2026-08-07T00:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/"},"modified":"2026-08-10T17:59:07","modified_gmt":"2026-08-10T17:59:07","slug":"diffusion-autoregressive-performance","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/","title":{"rendered":"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p>Giant Language Fashions (LLMs) have achieved state-of-the-art efficiency on a broad vary of Pure Language Processing (NLP) duties, together with doc processing and code era. Autoregressive Language Fashions (ARMs), which generate tokens sequentially conditioned on all earlier tokens, have been the predominant paradigm for LLMs. Whereas these fashions have achieved excessive accuracy throughout a spread of downstream duties, they exhibit low arithmetic depth as a result of inherent sequential dependency in next-token prediction. Not too long ago, Diffusion Language Fashions (DLMs) have emerged as a promising various structure. DLMs generate output tokens in parallel, mitigating the restrictions of sequential decoding. Nevertheless, the efficiency implications of DLMs relative to generally deployed ARMs should not totally understood. On this work, we current a complete research of the efficiency traits of ARMs and DLMs, combining theoretical evaluation with empirical profiling to characterize the trade-offs between these approaches. We present that though DLMs can obtain increased arithmetic depth than ARMs by leveraging parallelism throughout token positions, they fail to scale successfully with longer contexts. We then discover block-wise decoding for DLMs, which decouples arithmetic depth from sequence size and allows higher scaling to lengthy contexts (much like ARMs). We additionally study batched inference and discover that ARMs exhibit superior throughput as they profit extra from parallelism throughout sequences within the batch. Lastly, we spotlight alternatives for accelerating DLM inference, emphasizing that decreasing the variety of sampling steps is essential for open-source DLMs to realize decrease latency relative to ARMs.<\/p>\n<p>\u2020 Seoul Nationwide College\u2021 College of California, Berkeley\u00a7 ICSI\u00b6 LBNL\u2020\u2020 College of Texas at Austin* Advisory function<\/p><\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/machinelearning.apple.com\/research\/diffusion-autoregressive-performance\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Giant Language Fashions (LLMs) have achieved state-of-the-art efficiency on a broad vary of Pure Language Processing (NLP) duties, together with doc processing and code era. Autoregressive Language Fashions (ARMs), which generate tokens sequentially conditioned on all earlier tokens, have been the predominant paradigm for LLMs. Whereas these fashions have achieved excessive accuracy throughout a spread [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3562,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[2],"tags":[3920,3919,2389,50,293,3918,750,2002],"class_list":["post-3560","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-research-breakthroughs","tag-autoregressive","tag-characterization","tag-diffusion","tag-language","tag-models","tag-nexttoken","tag-performance","tag-prediction"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions - Future News 24<\/title>\n<meta name=\"description\" content=\"Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks\u2026\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks\u2026\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-07T00:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-10T17:59:07+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"1 minute\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions\",\"datePublished\":\"2026-08-07T00:00:00+00:00\",\"dateModified\":\"2026-08-10T17:59:07+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/\"},\"wordCount\":277,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"keywords\":[\"Autoregressive\",\"Characterization\",\"Diffusion\",\"Language\",\"Models\",\"NextToken\",\"performance\",\"Prediction\"],\"articleSection\":[\"AI Research &amp; Breakthroughs\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/\",\"name\":\"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"datePublished\":\"2026-08-07T00:00:00+00:00\",\"dateModified\":\"2026-08-10T17:59:07+00:00\",\"description\":\"Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks\u2026\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\",\"contentUrl\":\"https:\\\/\\\/mlr.cdn-apple.com\\\/media\\\/Home_1200x630_48225d82e9.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/07\\\/diffusion-autoregressive-performance\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions - Future News 24","description":"Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks\u2026","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/","og_locale":"en_US","og_type":"article","og_title":"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions - Future News 24","og_description":"Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks\u2026","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/","og_site_name":"Future News 24","article_published_time":"2026-08-07T00:00:00+00:00","article_modified_time":"2026-08-10T17:59:07+00:00","og_image":[{"url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"1 minute"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions","datePublished":"2026-08-07T00:00:00+00:00","dateModified":"2026-08-10T17:59:07+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/"},"wordCount":277,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","keywords":["Autoregressive","Characterization","Diffusion","Language","Models","NextToken","performance","Prediction"],"articleSection":["AI Research &amp; Breakthroughs"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/","name":"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/#primaryimage"},"thumbnailUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","datePublished":"2026-08-07T00:00:00+00:00","dateModified":"2026-08-10T17:59:07+00:00","description":"Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks\u2026","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/#primaryimage","url":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png","contentUrl":"https:\/\/mlr.cdn-apple.com\/media\/Home_1200x630_48225d82e9.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/07\/diffusion-autoregressive-performance\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Past Subsequent-Token Prediction: A Efficiency Characterization of Diffusion versus Autoregressive Language Fashions"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3560","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=3560"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3560\/revisions"}],"predecessor-version":[{"id":3561,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3560\/revisions\/3561"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/3562"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=3560"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=3560"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=3560"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}