{"id":3743,"date":"2026-08-14T11:08:00","date_gmt":"2026-08-14T11:08:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/"},"modified":"2026-08-14T18:59:13","modified_gmt":"2026-08-14T18:59:13","slug":"nvidia-nemotron-3-5-lightning","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/","title":{"rendered":"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Lengthy-running AI brokers typically spend most of their time on routine execution quite than troublesome reasoning. After making a plan, they might carry out a whole lot of device calls, file reads, validations, instructions, and formatting steps, so utilizing a frontier reasoning mannequin for each motion can turn into unnecessarily sluggish and costly.<\/p>\n<p>NVIDIA\u2019s Nemotron 3.5 Lightning takes a distinct strategy: a quick, environment friendly mannequin designed for high-volume agent execution. The concept is easy: use the costly mannequin to assume and the quick mannequin to work. On this article, we look at whether or not that structure can cut back price with out sacrificing agentic efficiency.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-is-nvidia-nemotron-3-5-lightning\">What&#8217;s NVIDIA Nemotron 3.5 Lightning?<\/h2>\n<p>Moreover, NVIDIA Nemotron 3.5 Lightning is an open-weight reasoning and instruction mannequin designed primarily for the execution layer of agentic programs.<\/p>\n<p>Its core specs are:<\/p>\n<div style=\"overflow-x:auto;margin:1em 0;\">SpecificationNemotron 3.5 LightningTotal parameters30BActive parameters3BArchitectureHybrid Mamba-2 + MoE + AttentionContext windowUp to 1M tokensInputTextOutputTextReasoningSupported and configurableTool callingSupportedQuantizationNVFP4, W4A16 optionsFull precision checkpointBF16Speculative decodingMTP, DSpark, DFlashRecommended temperature1.0Recommended top-p0.95LicenseOpenMDW 1.1Release dateAugust 11, 2026<\/div>\n<p>NVIDIA\u2019s official NVFP4 mannequin card additionally lists single-GPU deployment on a DGX Spark GB10 or H100, with help spanning Blackwell, Hopper and Ampere {hardware} relying on quantization.<\/p>\n<p>The mannequin is primarily supposed for English and programming languages, whereas Spanish, French, German, Italian and Japanese are additionally formally supported.<\/p>\n<p>That is necessary as a result of Nemotron 3.5 Lightning shouldn&#8217;t be evaluated as merely \u201cone other 30B mannequin.\u201d<\/p>\n<p>After all, its supposed job is way more particular.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-why-nvidia-built-an-execution-focused-model\">Why NVIDIA Constructed an Execution-Targeted Mannequin<\/h2>\n<p>Take into account a coding agent.<\/p>\n<p>It might first want to grasp a bug and develop a plan. That could be a troublesome reasoning downside.<\/p>\n<p>However after the plan exists, the agent might have to:<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"720\" height=\"940\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/agent-loop-flowchart.webp\" alt=\"Coding agent turn\" class=\"wp-image-256950\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/agent-loop-flowchart.webp 720w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/agent-loop-flowchart-230x300.webp 230w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/agent-loop-flowchart-150x196.webp 150w\" sizes=\"(max-width: 720px) 100vw, 720px\"\/><\/figure>\n<\/div>\n<p>Step one might deserve a frontier mannequin.<\/p>\n<p>Do all of the others?<\/p>\n<p>In all probability not.<\/p>\n<p>NVIDIA argues that long-running brokers spend a considerable portion of their workloads on precisely these high-volume execution operations, akin to device calls, validation and delegation. Utilizing a frontier reasoning mannequin for each execution step will increase each price and latency.<\/p>\n<p>Briefly, Nemotron 3.5 Lightning is NVIDIA\u2019s reply.<\/p>\n<p>A potential manufacturing structure turns into:<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter\"><img decoding=\"async\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/image-1-xhmgqa.webp\" alt=\"AI model request processing and task routing workflow\"\/><\/figure>\n<\/div>\n<p>Subsequent, this modifications how we must always take into consideration mannequin choice.<\/p>\n<p>As an alternative of asking:<\/p>\n<p>Lastly, which single mannequin ought to energy my agent?<\/p>\n<p>the extra helpful query turns into:<\/p>\n<p>Equally, which mannequin ought to deal with every kind of labor inside my agent?<\/p>\n<p>That&#8217;s the architectural thought behind Lightning.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-architecture-deep-dive\">Structure Deep Dive<\/h2>\n<p>In the meantime, Nemotron 3.5 Lightning makes use of one of many extra attention-grabbing architectures amongst present smaller agent fashions.<\/p>\n<p>NVIDIA describes it as a hybrid:<\/p>\n<p>Mamba-2<br \/>\n   +<br \/>\nCombination-of-Specialists<br \/>\n   +<br \/>\nSelective Consideration<br \/>\n   +<br \/>\nMulti-Token Prediction<\/p>\n<p>The mix issues as a result of every element solves a distinct effectivity downside.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-mixture-of-experts-30b-parameters-only-3b-active\">1. Combination-of-Specialists: 30B Parameters, Solely 3B Lively<\/h3>\n<p>Nemotron 3.5 Lightning incorporates roughly 30 billion whole parameters however prompts solely round 3 billion for every token.<\/p>\n<p>In a dense 30B mannequin, basically the entire community participates in inference.<\/p>\n<p>In an MoE mannequin:<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter\"><img decoding=\"async\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/image-2-xhmgqa-scaled.webp\" alt=\"Mixture of experts neural network architecture\"\/><\/figure>\n<\/div>\n<p>Then again, the router chooses solely a small subset of consultants.<\/p>\n<p>You due to this fact retain a lot of the representational capability of a bigger mannequin whereas doing computation nearer to a considerably smaller mannequin.<\/p>\n<p>That&#8217;s central to Lightning\u2019s throughput benefit.<\/p>\n<p>Printed runtime configuration additionally exposes 128 routed consultants plus a shared knowledgeable, with six routed consultants chosen per token. The configuration incorporates 52 hidden layers. Its hybrid layer sample resolves to Mamba, MoE and sparse Consideration parts quite than utilizing full self-attention at each layer. These are implementation-level configuration particulars, so builders ought to confirm them towards the precise checkpoint and runtime they deploy.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-2-mamba-2-layers\">2. Mamba-2 Layers<\/h3>\n<p>Conventional Transformers rely closely on consideration.<\/p>\n<p>Consideration is extraordinarily highly effective, however lengthy sequences turn into computationally costly.<\/p>\n<p>Though Mamba relies on state-space modeling and might course of sequences extra effectively.<\/p>\n<p>Nonetheless, Nemotron 3.5 Lightning doesn&#8217;t abandon consideration totally. As an alternative, NVIDIA makes use of Mamba-2 for a lot of the sequence processing whereas preserving chosen Consideration layers the place international token interplay stays helpful.<\/p>\n<p>Conceptually:<\/p>\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" width=\"720\" height=\"720\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/hybrid-block-stack.webp\" alt=\"Hybrid Block stack\" class=\"wp-image-256951\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/hybrid-block-stack.webp 720w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/hybrid-block-stack-300x300.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/hybrid-block-stack-150x150.webp 150w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/hybrid-block-stack-96x96.webp 96w\" sizes=\"(max-width: 720px) 100vw, 720px\"\/><\/figure>\n<p>This hybrid design is especially related for long-context brokers.<\/p>\n<p>As an alternative of paying full consideration prices all through your complete community, the mannequin mixes mechanisms optimized for various jobs.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-3-selective-attention\">3. Selective Consideration<\/h3>\n<p>Consideration continues to be necessary when tokens should instantly evaluate info throughout distant elements of the sequence.<\/p>\n<p>That issues for:<\/p>\n<p>lengthy paperwork<\/p>\n<p>source-code repositories<\/p>\n<p>multi-step device trajectories<\/p>\n<p>dialog historical past<\/p>\n<p>retrieved paperwork<\/p>\n<p>agent reminiscence<\/p>\n<p>Nemotron due to this fact retains chosen consideration layers as a substitute of switching to a pure state-space structure.<\/p>\n<p>The architectural philosophy is just not \u201cMamba as a substitute of Transformer.\u201d<\/p>\n<p>It&#8217;s use costly international consideration solely the place it provides enough worth.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-4-multi-token-prediction\">4. Multi-Token Prediction<\/h3>\n<p>Regular autoregressive LLMs study:<\/p>\n<p>Token 1 \u2192 predict Token 2Token 2 \u2192 predict Token 3Token 3 \u2192 predict Token 4<\/p>\n<p>Nemotron 3.5 Lightning contains Multi-Token Prediction, or MTP, layers that study to foretell a number of future tokens throughout coaching. Furthermore, NVIDIA added a devoted continued-pretraining stage for these MTP layers.<\/p>\n<p>MTP improves coaching alerts, however it additionally turns into helpful throughout inference.<\/p>\n<p>As an alternative of proposing solely:<\/p>\n<p>subsequent token<\/p>\n<p>the system can speculate about:<\/p>\n<p>token t+1token t+2token t+3&#8230;<\/p>\n<p>These candidates can then be verified effectively.<\/p>\n<p>Consequently, this is likely one of the mechanisms behind Lightning\u2019s excessive era throughput.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-why-is-nemotron-3-5-lightning-so-fast\">Why Is Nemotron 3.5 Lightning So Quick?<\/h2>\n<p>Then again, its velocity doesn&#8217;t come from one optimization. It&#8217;s the mixture of a number of.<\/p>\n<p>The most suitable choice due to this fact is determined by concurrency.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img decoding=\"async\" width=\"1500\" height=\"820\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/serving-path-guidance.webp\" alt=\"\" class=\"wp-image-256953\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/serving-path-guidance.webp 1500w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/serving-path-guidance-300x164.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/serving-path-guidance-768x420.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/serving-path-guidance-150x82.webp 150w\" sizes=\"(max-width: 1500px) 100vw, 1500px\"\/><\/figure>\n<\/div>\n<p>There isn&#8217;t any universally quickest configuration.<\/p>\n<p>MoE Sparsity: 30B parameters present capability, however solely about 3B are lively.<\/p>\n<p>Briefly, Hybrid Mamba Structure: Mamba reduces the necessity to carry out full consideration throughout each layer.<\/p>\n<p>Moreover, NVFP4 Quantization: Decrease-precision inference reduces reminiscence and compute necessities.<\/p>\n<p>As an alternative, Multi-Token Prediction: A number of future tokens could be proposed collectively.<\/p>\n<p>After all, Speculative Decoding: NVIDIA supplies three speculative approaches:<\/p>\n<p>MTP: Built-in instantly into the mannequin. NVIDIA recommends it notably for medium to excessive concurrency.<\/p>\n<p>Specifically, DSpark: A devoted draft mannequin optimized for DGX Spark and lower-concurrency data-center inference.<\/p>\n<p>Consequently, DFlash: A further draft mannequin that builders can benchmark towards MTP and DSpark for his or her workload.<\/p>\n<p>Whereas Nemotron 3.5 Lightning combines robust intelligence with as much as 4x output velocity of similar-sized fashions, putting it on the accuracy-speed Pareto frontier for high-volume agent workloads.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-nvidia-nemotron-3-5-lightning-benchmark-results\">NVIDIA Nemotron 3.5 Lightning Benchmark Outcomes<\/h2>\n<p>NVIDIA publishes each BF16 and NVFP4 outcomes throughout data, reasoning, coding, brokers, instruction following and lengthy context.<\/p>\n<p>In reality, the necessary statement is that quantization doesn&#8217;t dramatically collapse mannequin high quality.<\/p>\n<p>Listed here are the official reported outcomes. Benchmark-native items are preserved, so not each worth must be interpreted as a proportion.<\/p>\n<div style=\"overflow-x:auto;margin:1em 0;\">BenchmarkBF16NVFP4MMLU Pro81.9481.62AA-Omniscience17.5016.63GPQA Diamond, no tools75.4475.57HLE, text-only, no tools11.7210.47SciCode32.6031.38SWE-bench Verified51.5652.80SWE-bench Multilingual39.3336.47Terminal-Bench 2.124.5823.46PinchBench85.3783.43BrowseComp36.9736.81\u03c4\u00b3-bench Banking9.289.48GDPval-AA-V2832865IFBench loose71.8872.88AA-LCR52.0049.19<\/div>\n<p>Furthermore, NVIDIA says these evaluations had been run by a constant NeMo Fitness center and NeMo Evaluator-based harness and has printed benchmark recipes for reproducibility.<\/p>\n<p>In distinction, an attention-grabbing result&#8217;s how shut NVFP4 stays to BF16.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-nemotron-3-5-lightning-pricing\">Nemotron 3.5 Lightning Pricing<\/h2>\n<p>Nonetheless, pricing is barely extra sophisticated than a single quantity as a result of the mannequin is open-weight and obtainable by a number of routes.<\/p>\n<p>The next displays publicly listed pricing on August 12, 2026.<\/p>\n<div style=\"overflow-x:auto;margin:1em 0;\">\n<h3>Working vLLM regionally<\/h3>\n<p>Entry MethodCurrent CostContextBest ForNVIDIA Construct APIFree prototype endpoint1MTestingOpenRouter free routeFree1MQuick experimentationOpenRouter customary$0.05 enter \/ $0.20 output per 1M tokens262KSimple hosted APIFireworks serverlessSimilarly, $0.05 enter \/ $0.01 cached \/ $0.20 output per 1M262KProduction serverlessOllamaNo per-token mannequin feeRuntime dependentLocal\/personal useSelf-hosted vLLMInfrastructure costUp to 1MEnterprise\/self-hosting<\/p><\/div>\n<p>Subsequent, NVIDIA at the moment gives a free API endpoint for prototyping by construct.nvidia.com.<\/p>\n<p>In the meantime, OpenRouter lists each a free Nemotron 3.5 Lightning route with a 1M context and a normal route at the moment priced at $0.05 per million enter tokens and $0.20 per million output tokens. The usual OpenRouter route at the moment advertises a 262K context quite than the complete 1M mannequin functionality.<\/p>\n<p>Lastly, Fireworks at the moment lists precisely $0.05 per million enter tokens, $0.01 per million cached enter tokens and $0.20 per million output tokens, with a 262K serverless context window.<\/p>\n<p>Pricing and context limits can change shortly, notably in the course of the first weeks after a mannequin launch.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-how-to-access-nvidia-nemotron-3-5-lightning\">The right way to Entry NVIDIA Nemotron 3.5 Lightning<\/h2>\n<p>First, at launch, there are already a number of sensible methods to make use of the mannequin.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-option-1-nvidia-api\">Choice 1: NVIDIA API<\/h3>\n<p>Go to https:\/\/construct.nvidia.com\/ and login or join<\/p>\n<p>Click on in your profile image after which API keys.<\/p>\n<p>Generate a brand new API key.<\/p>\n<p>Now use this API for inference.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-option-2-ollama\">Choice 2: Ollama<\/h3>\n<p>Set up Ollama in your system from <\/p>\n<p>Run the next command in terminal to obtain and run Nemotron 3.5 lightening regionally.<\/p>\n<p>ollama run nemotron-3.5-lightning\u201d<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-option-3-openrouter\">Choice 3: OpenRouter<\/h3>\n<p>You may also use OpenRouter to run this mannequin. After all, its listed as a Free mannequin on OpenRouter. As an alternative, seize an API key and begin to use it<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-hands-on-using-nemotron-3-5-lightning-through-nvidia-api\">Palms-on: Utilizing Nemotron 3.5 Lightning By means of NVIDIA API<\/h2>\n<p>Nonetheless, NVIDIA exposes the mannequin by an OpenAI-compatible endpoint. The official instance makes use of nvidia\/nemotron-3.5-lightning-30b-a3b.<\/p>\n<p>Set up the shopper:<\/p>\n<p>pip set up openai<\/p>\n<p>Set your API key:<\/p>\n<p>export NVIDIA_API_KEY=&#8221;your_api_key&#8221;<\/p>\n<p>Now create a easy request:<\/p>\n<p>import os<br \/>\nfrom openai import OpenAI<\/p>\n<p>shopper = OpenAI(<br \/>\n    base_url=&#8221;https:\/\/combine.api.nvidia.com\/v1&#8243;,<br \/>\n    api_key=os.environ[&#8220;NVIDIA_API_KEY&#8221;]<br \/>\n)<\/p>\n<p>response = shopper.chat.completions.create(<br \/>\n    mannequin=&#8221;nvidia\/nemotron-3.5-lightning-30b-a3b&#8221;,<br \/>\n    messages=[<br \/>\n        {<br \/>\n            &#8220;role&#8221;: &#8220;user&#8221;,<br \/>\n            &#8220;content&#8221;: &#8220;&#8221;&#8221;<br \/>\n            A customer has submitted a warranty claim.<\/p>\n<p>            Purchase date: 2025-04-12<br \/>\n            Claim date: 2026-03-02<br \/>\n            Warranty duration: 12 months<br \/>\n            Damage type: manufacturing defect<\/p>\n<p>            Determine whether the claim is within the warranty period.<br \/>\n            Return JSON with:<br \/>\n            decision<br \/>\n            rationale<br \/>\n            &#8220;&#8221;&#8221;<br \/>\n        }<br \/>\n    ],<br \/>\n    temperature=1.0,<br \/>\n    top_p=0.95,<br \/>\n    max_tokens=2000,<br \/>\n    extra_body={<br \/>\n        &#8220;chat_template_kwargs&#8221;: {<br \/>\n            &#8220;enable_thinking&#8221;: True<br \/>\n        },<br \/>\n        &#8220;reasoning_budget&#8221;: 4000<br \/>\n    }<br \/>\n)<\/p>\n<p>print(response.decisions[0].message.content material)<\/p>\n<p>Output:<\/p>\n<p>{&#8220;determination&#8221;: &#8220;permitted&#8221;,&#8221;rationale&#8221;: &#8220;The guarantee interval begins on the acquisition date of 2025-04-12 and lasts for 12 months, ending on 2026-04-12. The declare was submitted on 2026-03-02, which falls inside the lively guarantee interval. Moreover, the harm is listed as a producing defect, which is often lined beneath customary guarantee phrases.&#8221;}<\/p>\n<p>This can be a higher first check than asking: Write a poem about AI.<\/p>\n<p>Nemotron 3.5 Lightning is designed for structured agent workloads, so check it accordingly.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>NVIDIA\u2019s primary argument is that future manufacturing AI programs might rely much less on a single big mannequin and extra on a coordinated structure of planners, routers, specialised staff, quick execution fashions, and verification layers. This represents a shift from maximizing mannequin dimension to optimizing how completely different fashions work collectively.<\/p>\n<p>In that structure, Nemotron 3.5 Lightning doesn&#8217;t have to be the neatest mannequin obtainable. Its worth comes from being environment friendly, quick, and succesful sufficient to deal with most routine agent duties whereas recognizing when more durable work must be escalated. NVIDIA is due to this fact optimizing for sensible, scalable agent execution quite than merely competing for the biggest or most clever mannequin.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-frequently-asked-questions\">Often Requested Questions<\/h2>\n<div class=\"schema-faq wp-block-yoast-faq-block\">\n<div class=\"schema-faq-section\" id=\"faq-question-1786691118615\">Q1. Is NVIDIA Nemotron 3.5 Lightning open supply? <\/p>\n<p class=\"schema-faq-answer\">A. NVIDIA supplies open mannequin weights, coaching knowledge, and recipes beneath the OpenMDW 1.1 license. It&#8217;s best described as an open-weight mannequin; please overview the governing license.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1786691118752\">Q2. How giant is Nemotron 3.5 Lightning? <\/p>\n<p class=\"schema-faq-answer\">A. It incorporates roughly 30B whole parameters whereas activating about 3B parameters per token.<\/p>\n<\/p><\/div>\n<div class=\"schema-faq-section\" id=\"faq-question-1786691118889\">Q3. What&#8217;s its context window? <\/p>\n<p class=\"schema-faq-answer\">A. The mannequin helps as much as 1 million tokens, though particular person suppliers can expose smaller limits.<\/p>\n<\/p><\/div><\/div>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n<p>                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_0fBqNLi.webp\" width=\"48\" height=\"48\" alt=\"Harsh Mishra\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p><\/div><\/div>\n<p>Harsh Mishra is an AI\/ML Engineer who spends extra time speaking to Massive Language Fashions than precise people. Obsessed with GenAI, NLP, and making machines smarter (so that they don\u2019t substitute him simply but). When not optimizing fashions, he\u2019s most likely optimizing his espresso consumption. \ud83d\ude80\u2615<\/p>\n<\/p><\/div><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to proceed studying and luxuriate in expert-curated content material.<\/h4>\n<p>                        Hold Studying for Free\n                    <\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2026\/08\/nvidia-nemotron-3-5-lightning\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Lengthy-running AI brokers typically spend most of their time on routine execution quite than troublesome reasoning. After making a plan, they might carry out a whole lot of device calls, file reads, validations, instructions, and formatting steps, so utilizing a frontier reasoning mannequin for each motion can turn into unnecessarily sluggish and costly. NVIDIA\u2019s Nemotron [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3745,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/Nemotro-3.5-Lightning.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[7],"tags":[15,417,4089,105,48,2805,445],"class_list":["post-3743","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-science-mlops","tag-agentic","tag-fast","tag-lightning","tag-model","tag-nemotron","tag-nvidias","tag-review"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin - Future News 24<\/title>\n<meta name=\"description\" content=\"A deep dive into NVIDIA&#039;s Nemotron 3.5 Lightning. Learn how this 30B MoE model uses Mamba-2 to deliver fast execution for AI agents.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin - Future News 24\" \/>\n<meta property=\"og:description\" content=\"A deep dive into NVIDIA&#039;s Nemotron 3.5 Lightning. Learn how this 30B MoE model uses Mamba-2 to deliver fast execution for AI agents.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-14T11:08:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-14T18:59:13+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/Nemotro-3.5-Lightning.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/Nemotro-3.5-Lightning.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"10 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin\",\"datePublished\":\"2026-08-14T11:08:00+00:00\",\"dateModified\":\"2026-08-14T18:59:13+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/\"},\"wordCount\":1986,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn.analyticsvidhya.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Nemotro-3.5-Lightning.webp\",\"keywords\":[\"Agentic\",\"Fast\",\"Lightning\",\"Model\",\"Nemotron\",\"Nvidias\",\"review\"],\"articleSection\":[\"Data Science &amp; MLOps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/\",\"name\":\"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/cdn.analyticsvidhya.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Nemotro-3.5-Lightning.webp\",\"datePublished\":\"2026-08-14T11:08:00+00:00\",\"dateModified\":\"2026-08-14T18:59:13+00:00\",\"description\":\"A deep dive into NVIDIA&#039;s Nemotron 3.5 Lightning. Learn how this 30B MoE model uses Mamba-2 to deliver fast execution for AI agents.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/#primaryimage\",\"url\":\"https:\\\/\\\/cdn.analyticsvidhya.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Nemotro-3.5-Lightning.webp\",\"contentUrl\":\"https:\\\/\\\/cdn.analyticsvidhya.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/Nemotro-3.5-Lightning.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/14\\\/nvidia-nemotron-3-5-lightning\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin - Future News 24","description":"A deep dive into NVIDIA&#039;s Nemotron 3.5 Lightning. Learn how this 30B MoE model uses Mamba-2 to deliver fast execution for AI agents.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/","og_locale":"en_US","og_type":"article","og_title":"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin - Future News 24","og_description":"A deep dive into NVIDIA&#039;s Nemotron 3.5 Lightning. Learn how this 30B MoE model uses Mamba-2 to deliver fast execution for AI agents.","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/","og_site_name":"Future News 24","article_published_time":"2026-08-14T11:08:00+00:00","article_modified_time":"2026-08-14T18:59:13+00:00","og_image":[{"url":"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/Nemotro-3.5-Lightning.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/Nemotro-3.5-Lightning.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"10 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin","datePublished":"2026-08-14T11:08:00+00:00","dateModified":"2026-08-14T18:59:13+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/"},"wordCount":1986,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/Nemotro-3.5-Lightning.webp","keywords":["Agentic","Fast","Lightning","Model","Nemotron","Nvidias","review"],"articleSection":["Data Science &amp; MLOps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/","name":"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/#primaryimage"},"thumbnailUrl":"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/Nemotro-3.5-Lightning.webp","datePublished":"2026-08-14T11:08:00+00:00","dateModified":"2026-08-14T18:59:13+00:00","description":"A deep dive into NVIDIA&#039;s Nemotron 3.5 Lightning. Learn how this 30B MoE model uses Mamba-2 to deliver fast execution for AI agents.","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/#primaryimage","url":"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/Nemotro-3.5-Lightning.webp","contentUrl":"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2026\/08\/Nemotro-3.5-Lightning.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/14\/nvidia-nemotron-3-5-lightning\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Nemotron 3.5 Lightning Evaluate: NVIDIA\u2019s Quick AI Agentic Mannequin"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3743","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=3743"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3743\/revisions"}],"predecessor-version":[{"id":3744,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3743\/revisions\/3744"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/3745"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=3743"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=3743"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=3743"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}