{"id":4448,"date":"2026-08-26T17:07:00","date_gmt":"2026-08-26T17:07:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/"},"modified":"2026-08-30T19:59:16","modified_gmt":"2026-08-30T19:59:16","slug":"experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/","title":{"rendered":"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">Alibaba launched the mannequin weights for Qwen3.8-Flash-Subsequent as a preview of the upcoming Qwen4 structure for builders to experiment with and consider. It\u2019s a multimodal mixture-of-experts (MoE) mannequin with a 125B-parameter important mannequin supplemented by an extra 51B N-gram embeddings, with 6B parameters activated per token. It has a local 262,144-token context window, extensible to 1M tokens with YaRN.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA supplies best-effort Day 0 practical help by SGLang, vLLM, and NVIDIA TensorRT LLM, validation throughout NVIDIA GB300 NVL72 for inference, and post-training recipes from NVIDIA NeMo AutoModel and NVIDIA NeMo\u00a0RL.\u00a0<\/p>\n<h2 id=\"architectural_innovations_for_long-context_inference\" class=\"wp-block-heading\">Architectural improvements for long-context inference<\/h2>\n<p class=\"wp-block-paragraph\">Qwen3.8-Flash-Subsequent is designed for high-volume, context-intensive purposes similar to agentic coding, doc processing, and tool-driven workflows. As context grows, consideration compute and KV cache reminiscence grow to be bottlenecks. The mannequin addresses each with a hybrid structure combining Gated DeltaNet (GDN) and Qwen Sparse Consideration (QSA). Three out of each 4 layers use GDN to repeatedly compress historic context right into a fixed-size recurrent state, eliminating KV cache progress as sequences lengthen. The\u00a0remaining layer makes use of QSA for exact retrieval throughout the total context.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Earlier sparse-attention approaches depend on token-level indexers that grow to be more and more computationally costly as context size grows. QSA aggregates the sequence into micro-blocks, estimates their significance on the block stage, and selects solely essentially the most related areas. This cuts consideration, compute, and indexing overhead inside every layer, making the design well-suited to architectures alternating between GDN and QSA\u00a0layers.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Alibaba\u2019s revealed benchmarks recommend that QSA can enhance the effectivity of 1M-token workloads. In contrast with full consideration, its consideration kernel delivered speedups of as much as 7.6x throughout prefill and 4.9x throughout decoding. In a cache-heavy on-line serving take a look at at a 1M-token context size and with a 90% prefix-cache hit charge, Qwen3.8-Flash-Subsequent achieved 8.6x the prefill throughput of Qwen3.7-Plus.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a948b7a89e25&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a948b7a89e25\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"1166\" height=\"1416\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1.webp\" alt=\"A diagram showing how GDN and QSA with MoE reduce memory and compute for large-context inference. \" class=\"wp-image-121865\" style=\"aspect-ratio:0.8234663398943096;width:802px;height:auto\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1.webp 1166w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-95x115.jpg 95w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-247x300.jpg 247w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-768x933.jpg 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-625x759.jpg 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-645x783.jpg 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-74x90.jpg 74w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-362x440.jpg 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-91x110.jpg 91w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-1024x1244.jpg 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-445x540.jpg 445w\" sizes=\"(max-width: 1166px) 100vw, 1166px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1166\" height=\"1416\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1.webp\" alt=\"A diagram showing how GDN and QSA with MoE reduce memory and compute for large-context inference. \" class=\"lazyload wp-image-121865\" style=\"aspect-ratio:0.8234663398943096;width:802px;height:auto\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1.webp 1166w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-95x115.jpg 95w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-247x300.jpg 247w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-768x933.jpg 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-625x759.jpg 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-645x783.jpg 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-74x90.jpg 74w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-362x440.jpg 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-91x110.jpg 91w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-1024x1244.jpg 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/qwen_arch-1-445x540.jpg 445w\" data-sizes=\"(max-width: 1166px) 100vw, 1166px\"\/><figcaption class=\"wp-element-caption\">Determine 1. Overview of Qwen3.8-Flash-Subsequent exhibiting three layers of GDN and one layer of QSA with MoE to scale back reminiscence and compute for large-context inference<\/figcaption><\/figure>\n<\/div>\n<figure class=\"wp-block-video aligncenter\"><figcaption class=\"wp-element-caption\">Video 1. Qwen3.8-Flash-Subsequent diagnoses and fixes a bug<\/figcaption><\/figure>\n<h2 id=\"running_qwen38-flash-next_on_nvidia_gb300_nvl72\" class=\"wp-block-heading\">Operating Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72<\/h2>\n<p class=\"wp-block-paragraph\">The GB300 NVL72 includes a rack-scale structure that integrates 72 NVIDIA Blackwell Extremely GPUs right into a single platform. Its massive, 72-GPU NVIDIA NVLink area permits environment friendly all-to-all communication at 130 TB\/s, eliminating bottlenecks that seem when professional visitors should cross conventional off-the-shelf networks. Operating on NVIDIA GB300 NVL72 delivers over 16K tokens per second per GPU and over 200 tokens per second per consumer, enabling builders to experiment with agentic coding purposes at excessive throughput and low latency.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a948b7a8ab3f&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a948b7a8ab3f\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"4000\" height=\"2300\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS.webp\" alt=\"A chart showing Qwen3.8-Flash-Next FP8 performance on NVIDIA GB300 NVL72 throughput vs. interactivity using TensorRT -LLM.\" class=\"wp-image-121860\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS.webp 4000w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-179x103.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-300x173.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-768x442.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-625x359.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-1536x883.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-2048x1178.png 2048w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-645x371.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-500x288.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-157x90.png 157w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-362x208.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-191x110.png 191w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-1024x589.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-939x540.png 939w\" sizes=\"(max-width: 4000px) 100vw, 4000px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"4000\" height=\"2300\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS.webp\" alt=\"A chart showing Qwen3.8-Flash-Next FP8 performance on NVIDIA GB300 NVL72 throughput vs. interactivity using TensorRT -LLM.\" class=\"lazyload wp-image-121860\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS.webp 4000w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-179x103.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-300x173.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-768x442.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-625x359.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-1536x883.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-2048x1178.png 2048w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-645x371.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-500x288.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-157x90.png 157w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-362x208.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-191x110.png 191w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-1024x589.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/16K-TPS-939x540.png 939w\" data-sizes=\"(max-width: 4000px) 100vw, 4000px\"\/><figcaption class=\"wp-element-caption\">Determine 2. A Pareto curve exhibiting Qwen3.8-Flash-Subsequent reaching peak throughput above 16K tokens per second per GPU on NVIDIA GB300 NVL72\u00a0<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Past rack-scale deployment, Qwen3.8-Flash-Subsequent additionally runs on native NVIDIA {hardware}, together with NVIDIA DGX Station, NVIDIA DGX Spark clusters, and workstations with 4 NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Version GPUs. Builders can prototype and consider agentic coding workflows on native {hardware} and scale the identical mannequin to GB300 NVL72 for manufacturing serving.\u00a0<\/p>\n<h2 id=\"post-train_qwen38-flash-next_and_serve_it_with_your_preferred_inference_engine\" class=\"wp-block-heading\">Put up-train Qwen3.8-Flash-Subsequent and serve it together with your most popular inference engine<\/h2>\n<p class=\"wp-block-paragraph\">Builders can fine-tune the mannequin for domain-specific use circumstances utilizing NVIDIA NeMo AutoModel, a PyTorch-native fine-tuning library with Day-0 Hugging Face checkpoint help. Practice straight on present checkpoints with out mannequin conversion, with help for full SFT or memory-efficient LoRA fine-tuning.\u00a0Customers can go a step to carry out reinforcement studying utilizing NVIDIA NeMo RL recipes.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA helps a number of inference stacks to fulfill a wide range of developer wants. SGLang, vLLM, and TokenSpeed present open-source inference recipes for builders requiring higher management over efficiency on the NVIDIA-accelerated platform. \u00a0<\/p>\n<h2 id=\"get_started_with_qwen38-flash-next\" class=\"wp-block-heading\">Get began with Qwen3.8-Flash-Subsequent<\/h2>\n<p class=\"wp-block-paragraph\">Check out the mannequin from QwenCloud. \u00a0<\/p>\n<p class=\"wp-block-paragraph\">Obtain the mannequin weights from Hugging Face or ModelScope.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Alibaba launched the mannequin weights for Qwen3.8-Flash-Subsequent as a preview of the upcoming Qwen4 structure for builders to experiment with and consider. It\u2019s a multimodal mixture-of-experts (MoE) mannequin with a 125B-parameter important mannequin supplemented by an extra 51B N-gram embeddings, with 6B parameters activated per token. It has a local 262,144-token context window, extensible to [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":4450,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/Agentic-AI-Qwen.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[15,1314,527,3162,81,2596,4551],"class_list":["post-4448","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-agentic","tag-coding","tag-experiment","tag-gb300","tag-nvidia","tag-nvl72","tag-qwen3-8flashnext"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding - Future News 24<\/title>\n<meta name=\"description\" content=\"Alibaba released the model weights for Qwen3.8&#x2d;Flash&#x2d;Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Alibaba released the model weights for Qwen3.8&#x2d;Flash&#x2d;Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-26T17:07:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-30T19:59:16+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/Agentic-AI-Qwen.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/Agentic-AI-Qwen.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"3 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding\",\"datePublished\":\"2026-08-26T17:07:00+00:00\",\"dateModified\":\"2026-08-30T19:59:16+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/\"},\"wordCount\":643,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Agentic-AI-Qwen.webp\",\"keywords\":[\"Agentic\",\"Coding\",\"Experiment\",\"GB300\",\"NVIDIA\",\"NVL72\",\"Qwen3.8FlashNext\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/\",\"name\":\"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Agentic-AI-Qwen.webp\",\"datePublished\":\"2026-08-26T17:07:00+00:00\",\"dateModified\":\"2026-08-30T19:59:16+00:00\",\"description\":\"Alibaba released the model weights for Qwen3.8&#x2d;Flash&#x2d;Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Agentic-AI-Qwen.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Agentic-AI-Qwen.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/26\\\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding - Future News 24","description":"Alibaba released the model weights for Qwen3.8&#x2d;Flash&#x2d;Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/","og_locale":"en_US","og_type":"article","og_title":"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding - Future News 24","og_description":"Alibaba released the model weights for Qwen3.8&#x2d;Flash&#x2d;Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/","og_site_name":"Future News 24","article_published_time":"2026-08-26T17:07:00+00:00","article_modified_time":"2026-08-30T19:59:16+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/Agentic-AI-Qwen.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/Agentic-AI-Qwen.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"3 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding","datePublished":"2026-08-26T17:07:00+00:00","dateModified":"2026-08-30T19:59:16+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/"},"wordCount":643,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/Agentic-AI-Qwen.webp","keywords":["Agentic","Coding","Experiment","GB300","NVIDIA","NVL72","Qwen3.8FlashNext"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/","name":"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/Agentic-AI-Qwen.webp","datePublished":"2026-08-26T17:07:00+00:00","dateModified":"2026-08-30T19:59:16+00:00","description":"Alibaba released the model weights for Qwen3.8&#x2d;Flash&#x2d;Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate.","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/Agentic-AI-Qwen.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/Agentic-AI-Qwen.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/26\/experiment-with-qwen3-8-flash-next-on-nvidia-gb300-nvl72-for-agentic-coding\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Experiment with Qwen3.8-Flash-Subsequent on NVIDIA GB300 NVL72 for Agentic Coding"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4448","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=4448"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4448\/revisions"}],"predecessor-version":[{"id":4449,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4448\/revisions\/4449"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/4450"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=4448"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=4448"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=4448"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}