{"id":3788,"date":"2026-08-11T13:01:00","date_gmt":"2026-08-11T13:01:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/"},"modified":"2026-08-15T17:59:05","modified_gmt":"2026-08-15T17:59:05","slug":"nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/","title":{"rendered":"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">Lengthy-running AI brokers spend most of their time on high-volume execution: instrument calls, end result validation, and subagent delegation. Utilizing a frontier reasoning mannequin for each execution step provides price and latency.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) mannequin with 3B energetic parameters constructed for that execution layer of always-on brokers. It&#8217;s designed for harnesses like OpenClaw and Hermes Agent\u2014all supported by the NVIDIA NemoClaw open supply safety and administration stack for working always-on AI brokers.<\/p>\n<p class=\"wp-block-paragraph\">The NVIDIA Nemotron open mannequin household is sort of a software program library, with every launch constantly enhancing accuracy and pace. As these fashions evolve, the speedy maturation of mannequin routing and orchestration can be underway.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">That is vital as a result of builders more and more construct purposes with a system of fashions. Frontier reasoning fashions similar to Nemotron 3 Extremely deal with orchestration and sophisticated planning whereas smaller, extra environment friendly fashions deal with the high-volume execution layer.<\/p>\n<p class=\"wp-block-paragraph\">This put up introduces NVIDIA Nemotron 3.5 Lightning and explains how its smaller MoE design is optimized for high-volume, low-latency execution in autonomous brokers. It additionally particulars the inference and coaching improvements that energy it. Lastly, the put up additionally introduces NVIDIA NeMo Switchyard, a library that intelligently routes every process to the very best mannequin for the job.<\/p>\n<h2 id=\"why_is_nemotron_35_lightning_ideal_for_long-running_ai_agents\u00a0\" class=\"wp-block-heading\">Why is Nemotron 3.5 Lightning excellent for long-running AI brokers?\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning is a customizable open 30B MoE mannequin with 3B energetic parameters, offering optimum high-volume execution for autonomous brokers. MoEs are quick and environment friendly as a result of a router sends every token to just some of its many specialists, so solely a fraction of the mannequin\u2019s parameters run per token. This offers the capability of a bigger dense mannequin on the compute price of a small one.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning is the smallest member of the Nemotron 3 mannequin household and ships with most of the identical methods confirmed throughout the household, together with:<\/p>\n<p>Speculative decoding: Multi-token prediction was included throughout Nemotron 3.5 Lightning coaching (as for Nemotron 3 Tremendous and Nemotron 3 Extremely). Nemotron 3.5 Lightning additionally ships with DFlash and DSpark, enabling extra complete inference optimization throughout a spread of serving eventualities.<\/p>\n<p>Harness-optimized coaching: The mannequin is skilled for well-liked agent harnesses, enabling brokers to make extra correct calls whereas decreasing latency for high-volume duties.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The result&#8217;s a mannequin constructed for execution-focused, excessive name volumes, and low latency\u2014all at a measurement that deploys wherever from an NVIDIA DGX Spark to a knowledge middle.<\/p>\n<figure class=\"wp-block-embed aligncenter is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio\">\n<p>\n<span class=\"embed-youtube\" style=\"text-align:center; display: block;\"><\/span>\n<\/p><figcaption class=\"wp-element-caption\">Video 1. Learn to deploy NVIDIA Nemotron 3.5 Lightning on DGX Spark and use it for quick, high-volume agentic workloads<\/figcaption><\/figure>\n<h2 id=\"customize_nemotron_35_lightning_out_of_the_box\u00a0\" class=\"wp-block-heading\">Customise Nemotron 3.5 Lightning out of the field\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Fashions earn their place in specialised AI agent techniques once they\u2019re tailored to the job. And Lightning-class fashions are extremely customizable: small fashions fine-tune quicker, cheaper, and on much more modest {hardware} than their bigger counterparts.<\/p>\n<p class=\"wp-block-paragraph\">You&#8217;ll be able to customise Nemotron 3.5 Lightning out of the field to suit your workload. As with each Nemotron open mannequin launch, the weights, coaching information, and recipes are launched as permissively as doable underneath OpenMDW-1.1, so you may:<\/p>\n<p class=\"wp-block-paragraph\">This launch contains Nemotron-RL Agentic Terminal Pivot, an open agentic reinforcement studying dataset used to coach among the coding agent capabilities.\u00a0<\/p>\n<h2 id=\"route_work_to_the_right_model_using_nemo_switchyard\" class=\"wp-block-heading\">Route work to the fitting mannequin utilizing NeMo Switchyard<\/h2>\n<p class=\"wp-block-paragraph\">Whereas frontier fashions might win the headlines, fashions like Nemotron 3.5 Lightning earn their medals within the trenches. They deal with requests like git pull, validate instrument outputs, format outcomes, and run the routine calls that dominate any long-running agent\u2019s token funds.<\/p>\n<p class=\"wp-block-paragraph\">Mannequin routing and orchestration assist make this division of labor extra accessible. They&#8217;re now\u00a0 accessible by means of NVIDIA NeMo Switchyard. Switchyard can expose Nemotron 3.5 Lightning as a routing goal alongside your open and closed fashions, so each request lands on essentially the most succesful and environment friendly mannequin that may deal with it. Plans route as much as the frontier, execution routes right down to Lightning, making certain that your tokens are spent effectively and successfully.<\/p>\n<h2 id=\"how_does_nemotron_35_lightning_perform_on_the_accuracy-speed_pareto_frontier\u00a0\" class=\"wp-block-heading\">How does Nemotron 3.5 Lightning carry out on the accuracy-speed Pareto frontier?\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning delivers main accuracy on the highest output pace in its class, profitable the accuracy-versus-speed Pareto frontier on the Synthetic Evaluation Intelligence Index. This index combines 9 evaluations to measure mannequin efficiency throughout agentic duties, coding, scientific reasoning, and normal intelligence.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning combines sturdy intelligence with as much as 4x output pace of similar-sized fashions, inserting it on the accuracy-speed Pareto frontier for high-volume agent workloads.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a80a8e846094&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a80a8e846094\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"2582\" height=\"1280\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning.webp\" alt=\"Artificial Analysis Intelligence Index versus output speed scatter, with Nemotron 3.5 Lightning in the winning quadrant.&#10;\" class=\"wp-image-121103\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning.webp 2582w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-179x89.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-300x149.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-768x381.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-625x310.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-1536x761.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-2048x1015.png 2048w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-645x320.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-500x248.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-160x79.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-362x179.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-222x110.png 222w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-1024x508.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-960x476.png 960w\" sizes=\"(max-width: 2582px) 100vw, 2582px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"2582\" height=\"1280\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning.webp\" alt=\"Artificial Analysis Intelligence Index versus output speed scatter, with Nemotron 3.5 Lightning in the winning quadrant.&#10;\" class=\"lazyload wp-image-121103\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning.webp 2582w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-179x89.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-300x149.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-768x381.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-625x310.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-1536x761.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-2048x1015.png 2048w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-645x320.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-500x248.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-160x79.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-362x179.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-222x110.png 222w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-1024x508.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/artificial-analylsis-intelligence-index-versus-output-speed-nemotron-3.5-lightning-960x476.png 960w\" data-sizes=\"(max-width: 2582px) 100vw, 2582px\"\/><figcaption class=\"wp-element-caption\">Determine 1. Nemotron 3.5 Lightning defines the accuracy-speed Pareto frontier for small open fashions on the Synthetic Evaluation Intelligence Index leaderboard<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Agent effectivity in the end comes right down to how rapidly a mannequin completes helpful work and never merely how briskly it generates tokens. On PinchBench, Nemotron 3.5 Lightning reaches 86% accuracy whereas finishing 10,000 duties 30% quicker than Qwen3.6 35B at related accuracy.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Increased inference throughput and token effectivity locations Nemotron 3.5 Lightning on the effectivity frontier, serving to always-on brokers end high-volume work quicker.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a80a8e8470e1&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a80a8e8470e1\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"944\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks.webp\" alt=\"Chart comparing PinchBench accuracy with time to complete 10,000 tasks. Nemotron 3.5 Lightning reaches similar accuracy as Qwen3.6 35B 30% faster. &#10;\" class=\"wp-image-121104\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-179x85.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-300x142.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-768x363.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-625x295.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-1536x725.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-645x305.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-500x236.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-160x76.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-362x171.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-233x110.png 233w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-1024x484.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-960x453.png 960w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"944\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks.webp\" alt=\"Chart comparing PinchBench accuracy with time to complete 10,000 tasks. Nemotron 3.5 Lightning reaches similar accuracy as Qwen3.6 35B 30% faster. &#10;\" class=\"lazyload wp-image-121104\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-179x85.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-300x142.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-768x363.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-625x295.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-1536x725.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-645x305.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-500x236.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-160x76.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-362x171.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-233x110.png 233w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-1024x484.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nemotron-3.5-lightning-efficiency-frontier-agentic-tasks-960x453.png 960w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 2. Nemotron 3.5 Lightning leads the effectivity frontier by finishing agentic duties as much as 30% quicker at comparable accuracies\u00a0<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"how_does_nemotron_35_lightning_deliver_speed_without_compromising_accuracy\" class=\"wp-block-heading\">How does Nemotron 3.5 Lightning ship pace with out compromising accuracy?<\/h2>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning delivers pace and customization with out compromising accuracy by means of speculative decoding, and quantization.<\/p>\n<h3 id=\"speculative_decoding\u00a0\" class=\"wp-block-heading\">Speculative decoding\u00a0<\/h3>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning is constructed to rapidly generate tokens and has the power to generate a number of tokens by means of speculative decoding. It is a course of whereby the mannequin, or draft mannequin, will draft some variety of tokens that are effectively reviewed. Nemotron 3.5 Lightning underwent a devoted pretraining stage to bake multi-token prediction (MTP) into the mannequin, as with Nemotron 3 Tremendous and Extremely. After coaching, a devoted MTP-boosting part additional improved MTP accuracy.<\/p>\n<p class=\"wp-block-paragraph\">Past MTP, two draft fashions are supplied with Nemotron 3.5 Lightning: DSpark, which is advisable for DGX Spark inference workloads and low concurrency information middle workloads. MTP is finest fitted to medium to excessive concurrency, with the optimum draft size lowering as concurrency will increase. NVIDIA can be releasing a DFlash draft mannequin, which will be measured towards the others and will carry out finest on your workloads.\u00a0<\/p>\n<h3 id=\"quantization\u00a0\" class=\"wp-block-heading\">Quantization\u00a0<\/h3>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning ships with an NVFP4 checkpoint alongside BF16, utilizing the identical specialised NVFP4 kernels that energy Nemotron 3 Extremely throughout NVIDIA Blackwell, NVIDIA Hopper, and NVIDIA Ampere GPUs. The identical file serves simply as effectively in information facilities because it does in your desktop DGX Spark.<\/p>\n<h2 id=\"how_is_nemotron_35_lightning_ideal_for_local_ai\" class=\"wp-block-heading\">How is Nemotron 3.5 Lightning excellent for native AI?<\/h2>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning makes succesful agentic AI accessible on native techniques together with NVIDIA Jetson, GeForce RTX 5090, and DGX Spark.\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA has labored with various groups together with EXO Labs to grasp how this mannequin performs on DGX Spark.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a80a8e848304&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a80a8e848304\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"3134\" height=\"1585\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1.webp\" alt=\"A scatter\/line chart showing the DGX Spark performance \u201cIntelligence \/ Speed Frontier,\u201d plotting utilization over task time. It compares several model configurations along a Pareto frontier curve, with callouts indicating parameter counts and results.&#10;\" class=\"wp-image-121119\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1.webp 3134w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-179x91.jpg 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-300x152.jpg 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-768x388.jpg 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-625x316.jpg 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-1536x777.jpg 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-2048x1036.jpg 2048w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-645x326.jpg 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-500x253.jpg 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-160x81.jpg 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-362x183.jpg 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-218x110.jpg 218w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-1024x518.jpg 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-960x486.jpg 960w\" sizes=\"(max-width: 3134px) 100vw, 3134px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"3134\" height=\"1585\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1.webp\" alt=\"A scatter\/line chart showing the DGX Spark performance \u201cIntelligence \/ Speed Frontier,\u201d plotting utilization over task time. It compares several model configurations along a Pareto frontier curve, with callouts indicating parameter counts and results.&#10;\" class=\"lazyload wp-image-121119\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1.webp 3134w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-179x91.jpg 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-300x152.jpg 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-768x388.jpg 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-625x316.jpg 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-1536x777.jpg 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-2048x1036.jpg 2048w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-645x326.jpg 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-500x253.jpg 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-160x81.jpg 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-362x183.jpg 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-218x110.jpg 218w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-1024x518.jpg 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/local-ai-dgx-spark-frontier-with-nemotron-3.5-lightning-1-960x486.jpg 960w\" data-sizes=\"(max-width: 3134px) 100vw, 3134px\"\/><figcaption class=\"wp-element-caption\">Determine 3. Nemotron 3.5 Lightning sits proper on the Pareto frontier for small open fashions on the EXO Labs native.ai leaderboard<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">As well as, you may run Nemotron 3.5 Lightning throughout the business customary set of instruments, together with LM Studio, llama.cpp, Ollama, and Unsloth.\u00a0<\/p>\n<h2 id=\"partner_ecosystem\" class=\"wp-block-heading\">Companion ecosystem<\/h2>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning is supported by a rising ecosystem of companions throughout harnesses, customization, deployment, and inference, together with:<\/p>\n<p>Submit-training: AgileRL, Utilized Compute, Deep Cogito, distil labs, Fastino Labs, Locai Labs, Prime Mind, Cheap, Considering Machines Lab, Thoughtworks, Trajectory, Uniphore<\/p>\n<p>Inference software program: Ollama, Exo, Canonical, LM Studio, Unsloth<\/p>\n<p>Harnesses and agent frameworks: Aible, Cline, Manufacturing unit AI, Hermes Agent, Kilo Code, LangChain, LM Studio Bionic, OpenClaw, OpenCode, OpenHands, Pi<\/p>\n<p>Cloud service supplier platforms: Amazon SageMaker JumpStart, Google Cloud Gemini Enterprise Agent Platform, MSFT Foundry, OCI Enterprise AI<\/p>\n<p>GSI: Accenture, Tata Consultancy Providers, Tech Mahindra, Wipro<\/p>\n<p>AI natives: Arcos Labs, CodeRabbit, Dream, Harvey<\/p>\n<p>Hosted inference service suppliers: Baseten, Bitdeer AI, BlackBox AI, CoreWeave, Crusoe, DeepInfra, Fireworks AI, FriendliAI, GMI Cloud, Modal, Nebius, Collectively AI<\/p>\n<h2 id=\"start_building_with_nemotron_35_lightning\" class=\"wp-block-heading\">Begin constructing with Nemotron 3.5 Lightning<\/h2>\n<p class=\"wp-block-paragraph\">Nemotron 3.5 Lightning is absolutely open\u2014weights, information, and recipes\u2014so you may adapt it to your workflows and deploy it wherever. To get began, attempt it on construct.nvidia.com or by means of OpenRouter. Obtain the weights from Hugging Face, and ModelScope. Wish to dive deeper?<\/p>\n<p class=\"wp-block-paragraph\">Keep updated on NVIDIA Nemotron by subscribing to NVIDIA information and following NVIDIA AI on LinkedIn, X, Discord, and YouTube.<\/p>\n<p class=\"wp-block-paragraph\">Go to the Nemotron developer web page for sources to get began. Discover open Nemotron fashions and datasets on Hugging Face, ModelScope, and Blueprints on construct.nvidia.com.<\/p>\n<p class=\"wp-block-paragraph\">Have interaction with Nemotron livestreams, tutorials, and the developer neighborhood on the NVIDIA discussion board and Discord.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Lengthy-running AI brokers spend most of their time on high-volume execution: instrument calls, end result validation, and subagent delegation. Utilizing a frontier reasoning mannequin for each execution step provides price and latency.\u00a0 NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) mannequin with 3B energetic parameters constructed for that execution layer of always-on brokers. [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3790,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/agentic-ai-nemotron-3.5-lightning.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[3569,210,3004,3980,417,4089,209,48,81,1304,1830],"class_list":["post-3788","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-accurate","tag-agents","tag-delivers","tag-execution","tag-fast","tag-lightning","tag-longrunning","tag-nemotron","tag-nvidia","tag-specialized","tag-task"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers - Future News 24<\/title>\n<meta name=\"description\" content=\"Long&#x2d;running AI agents spend most of their time on high&#x2d;volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning&#8230;\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Long&#x2d;running AI agents spend most of their time on high&#x2d;volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning&#8230;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-11T13:01:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-15T17:59:05+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/agentic-ai-nemotron-3.5-lightning.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/agentic-ai-nemotron-3.5-lightning.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"7 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers\",\"datePublished\":\"2026-08-11T13:01:00+00:00\",\"dateModified\":\"2026-08-15T17:59:05+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/\"},\"wordCount\":1406,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/agentic-ai-nemotron-3.5-lightning.webp\",\"keywords\":[\"accurate\",\"Agents\",\"delivers\",\"Execution\",\"Fast\",\"Lightning\",\"LongRunning\",\"Nemotron\",\"NVIDIA\",\"Specialized\",\"Task\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/\",\"name\":\"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/agentic-ai-nemotron-3.5-lightning.webp\",\"datePublished\":\"2026-08-11T13:01:00+00:00\",\"dateModified\":\"2026-08-15T17:59:05+00:00\",\"description\":\"Long&#x2d;running AI agents spend most of their time on high&#x2d;volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning&#8230;\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/agentic-ai-nemotron-3.5-lightning.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/agentic-ai-nemotron-3.5-lightning.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/11\\\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers - Future News 24","description":"Long&#x2d;running AI agents spend most of their time on high&#x2d;volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning&#8230;","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/","og_locale":"en_US","og_type":"article","og_title":"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers - Future News 24","og_description":"Long&#x2d;running AI agents spend most of their time on high&#x2d;volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning&#8230;","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/","og_site_name":"Future News 24","article_published_time":"2026-08-11T13:01:00+00:00","article_modified_time":"2026-08-15T17:59:05+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/agentic-ai-nemotron-3.5-lightning.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/agentic-ai-nemotron-3.5-lightning.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"7 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers","datePublished":"2026-08-11T13:01:00+00:00","dateModified":"2026-08-15T17:59:05+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/"},"wordCount":1406,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/agentic-ai-nemotron-3.5-lightning.webp","keywords":["accurate","Agents","delivers","Execution","Fast","Lightning","LongRunning","Nemotron","NVIDIA","Specialized","Task"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/","name":"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/agentic-ai-nemotron-3.5-lightning.webp","datePublished":"2026-08-11T13:01:00+00:00","dateModified":"2026-08-15T17:59:05+00:00","description":"Long&#x2d;running AI agents spend most of their time on high&#x2d;volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning&#8230;","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/agentic-ai-nemotron-3.5-lightning.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/agentic-ai-nemotron-3.5-lightning.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/11\/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"NVIDIA Nemotron 3.5 Lightning Delivers Quick, Correct Specialised Activity Execution for Lengthy-Operating Brokers"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3788","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=3788"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3788\/revisions"}],"predecessor-version":[{"id":3789,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3788\/revisions\/3789"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/3790"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=3788"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=3788"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=3788"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}