{"id":3521,"date":"2026-07-30T16:00:00","date_gmt":"2026-07-30T16:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/"},"modified":"2026-08-09T21:59:06","modified_gmt":"2026-08-09T21:59:06","slug":"nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/","title":{"rendered":"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">Two AI computing clusters constructed from equivalent NVIDIA H100, GB200 NVL72, or GB300 NVL72 methods can ship materially completely different coaching throughput. We routinely see 8% to 12% gaps between accomplice deployments and the corresponding NVIDIA reference structure (RA) on the identical workload, similar mannequin, similar world batch measurement.<\/p>\n<p class=\"wp-block-paragraph\">The trigger is commonly a stack of configuration selections within the kernel, hypervisor, BIOS, and NVIDIA Collective Communications Library (NCCL) settings, every costing a number of %, that compound into a niche massive sufficient to overlook the 95% threshold required for NVIDIA Exemplar Cloud validation.<\/p>\n<p class=\"wp-block-paragraph\">This publish walks by way of 4 debugging investigations from actual accomplice clusters. Every diagnostic isolates a definite layer of the stack: system reminiscence administration unit (SMMU) and page-table habits on NVIDIA Grace CPU; energy administration and non-uniform reminiscence entry (NUMA) placement on x86-based CPU; NVIDIA NCCL queue-pair concurrency on 1.6 Tbps materials; and silent hardware-installation defects. The publish additionally reveals the precise sign in perf, NVIDIA Nsight Programs, or NVIDIA NCCL checks that pointed to the basis trigger, alongside the tuning change that closed the hole.<\/p>\n<p class=\"wp-block-paragraph\">Infrastructure engineers and efficiency architects who already run these benchmarks can profit from these diagnostic patterns we use internally to run towards their very own clusters earlier than formal RA validation.<\/p>\n<h3 id=\"prerequisites\" class=\"wp-block-heading\">Stipulations<\/h3>\n<p class=\"wp-block-paragraph\">To breed the diagnostics on this publish, you will want:<\/p>\n<p>An NVIDIA HGX H100, HGX H200, HGX B200, GB200 NVL72, or GB300 NVL72 methods cluster with NVIDIA Quantum InfiniBand or RoCE interconnect.<\/p>\n<p>A distributed coaching workload with steady iteration timing\u2014NVIDIA NeMo on Llama 3 mannequin, NVIDIA Nemotron, or DeepSeek configuration is an inexpensive reference.<\/p>\n<p>Root entry on no less than one node for perf, BIOS\/UEFI adjustments, and kernel parameter adjustments.<\/p>\n<p>nccl-tests constructed towards the identical NCCL model your coaching stack makes use of, NVIDIA Nsight Programs, and Linux perf with kernel symbols obtainable.<\/p>\n<h2 id=\"common_patterns_behind_training_performance_gaps\" class=\"wp-block-heading\">Widespread patterns behind coaching efficiency gaps<\/h2>\n<p class=\"wp-block-paragraph\">Latest Exemplar coaching engagements present that efficiency gaps not often come from a single apparent failure. Extra typically, they arrive from configuration particulars that grow to be seen solely underneath workload strain. Some recurring patterns embrace:<\/p>\n<p>Grace and virtualization readiness: Lacking platform capabilities, SMMU overhead, IOMMU habits, or page-size settings that don\u2019t match the anticipated configuration.<\/p>\n<p>CPU energy and course of placement: Cores operating under anticipated turbo frequency, ranks or helper threads positioned on the flawed cores, or NUMA\/PCT bindings that don\u2019t match the platform topology.<\/p>\n<p>Runtime topology: Host topology information or NCCL settings which can be appropriate on the node however lacking contained in the workload container or launcher atmosphere.<\/p>\n<p>Cloth and collective habits: NCCL settings that don\u2019t match the goal material, message measurement, or scale of the coaching workload.<\/p>\n<p>Software-to-platform binding: Coaching processes binding by core ID or rank order as an alternative of topology-aware affinity.<\/p>\n<p class=\"wp-block-paragraph\">These aren\u2019t the one causes of coaching efficiency gaps, and checking them doesn\u2019t exchange validation with actual purposes.<\/p>\n<p class=\"wp-block-paragraph\">4 case research under present how these patterns appeared in current coaching work: what sign uncovered the difficulty, what modified, and the way the repair was verified. The order isn\u2019t a common triage sequence; the correct place to begin relies on the workload, platform, and first profiler sign.<\/p>\n<p class=\"wp-block-paragraph\">Layer: Virtualization and SMMU<\/p>\n<p class=\"wp-block-paragraph\">A GB200 NVL72 accomplice deployment operating DeepSeek-V3 Combination-of-Consultants (MoE) FP8 pre-training inside a VM was producing iteration occasions 12% to 14% longer than the bare-metal RA. Pre-training recipes for dense fashions like Llama 3 70B ran inside 3% of RA efficiency, whereas DeepSeek-V3 MoE, which points many small kernels per iteration, was the outlier.<\/p>\n<p class=\"wp-block-paragraph\">Nsight Programs traces captured on the accomplice cluster confirmed considerably increased CPU overhead for tiny kernel areas of the workload. Microbenchmarks focusing on simply the CPU single thread efficiency demonstrated close to equivalent efficiency on the accomplice and RA cluster nodes. This indicated {that a} 30-second perf document -a -g seize on the host, considered with perf report, surfaced an sudden prime body: 24% of CPU cycles spent on arm_smmu_cmdq_issue_cmdlist.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a78f80b127c3&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a78f80b127c3\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1200\" height=\"831\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3.webp\" alt=\"Linux perf icicle graph showing arm_smmu_cmdq_issue_cmdlist consuming ~24% of CPU time via memory-unmapping and TLB-flush paths during a virtualized DeepSeek-V3 workload, indicating significant Arm SMMU command-queue overhead.&#10;\" class=\"wp-image-120321\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3.webp 1200w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-166x115.png 166w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-300x208.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-768x532.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-625x433.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-645x447.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-433x300.png 433w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-130x90.png 130w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-362x251.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-159x110.png 159w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-1024x709.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-780x540.png 780w\" sizes=\"(max-width: 1200px) 100vw, 1200px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1200\" height=\"831\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3.webp\" alt=\"Linux perf icicle graph showing arm_smmu_cmdq_issue_cmdlist consuming ~24% of CPU time via memory-unmapping and TLB-flush paths during a virtualized DeepSeek-V3 workload, indicating significant Arm SMMU command-queue overhead.&#10;\" class=\"lazyload wp-image-120321\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3.webp 1200w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-166x115.png 166w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-300x208.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-768x532.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-625x433.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-645x447.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-433x300.png 433w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-130x90.png 130w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-362x251.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-159x110.png 159w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-1024x709.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image7-3-780x540.png 780w\" data-sizes=\"(max-width: 1200px) 100vw, 1200px\"\/><figcaption class=\"wp-element-caption\">Determine 1. Linux perf icicle graph highlighting Arm SMMU command-queue invalidation overhead throughout DeepSeek-V3 FP8 pre-training on a virtualized NVIDIA GB200 NVL72 system<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">arm_smmu_cmdq_issue_cmdlist is the operate that submits invalidation instructions to the Arm SMMU\u2019s command queue. Underneath virtualization, each map\/unmap leading to visitor invalidation traps the host and serializes by way of a single command queue, producing the spinlock competition seen within the profile. Digital Command Queue (VCMDQ) is a characteristic obtainable by way of the Command Queue Virtualization extension to the usual Arm SMMUv3 and permits the visitor to subject SMMU invalidation instructions on to {hardware} with out VM exits.<\/p>\n<p class=\"wp-block-paragraph\">The repair: Allow CMDQV\/VCMDQ within the host kernel on the accomplice cluster and expose it to the visitor. This requires a kernel constructed with the tegra241-cmdqv driver and the corresponding hypervisor assist; current QEMU\/libvirt variations have added a cmdqv IOMMU attribute to show it to visitors.<\/p>\n<p class=\"wp-block-paragraph\">After this modification, linux perf confirmed arm_smmu_cmdq_issue_cmdlist falling out of the highest frames and dTLB miss charges returning to bare-metal parity. The MoE iteration-time hole narrowed to inside RA tolerance from 12%.<\/p>\n<p class=\"wp-block-paragraph\">The takeaway is that Grace-based virtualized deployments want the VM stack to show the correct SMMU capabilities for memory-mapping-heavy workloads. With CMDQV\/VCMDQ enabled within the host kernel and uncovered to the visitor, the platform can keep away from pointless SMMU serialization and return MoE coaching efficiency to inside RA tolerance.<\/p>\n<p class=\"wp-block-paragraph\">The subsequent layer down is the CPU itself, the place the failure mode seems utterly completely different.<\/p>\n<h2 id=\"case_study_2_h100_cluster_losing_12%_to_cpu_contention_and_numa_misbinding\" class=\"wp-block-heading\">Case examine 2: H100 cluster shedding 12% to CPU competition and NUMA misbinding<\/h2>\n<p class=\"wp-block-paragraph\">Layer: CPU energy and course of placement.<\/p>\n<p class=\"wp-block-paragraph\">A accomplice\u2019s H100 SXM5 cluster, operating the identical NCCL model and NeMo container as NVIDIA\u2019s HGX RA, was operating Llama 3 70B pre-training 12% slower than reference. Not like the GB200 NVL72 case, this wasn\u2019t a kernel-level subject; every part occurred in consumer house and BIOS.<\/p>\n<p class=\"wp-block-paragraph\">Two issues stood out:<\/p>\n<p>CPU frequency: turbostat -i 1 throughout coaching confirmed busy cores pegged at 3.0 GHz, regardless of the SKU being rated for 3.8 GHz turbo. Idle cores had been additionally at 3.0 GHz, with C-states sitting in C1 slightly than dropping to C6.<\/p>\n<p>NUMA-remote site visitors: numastat -p  confirmed roughly 18% of the coaching course of\u2019s reminiscence accesses going to the distant NUMA node<\/p>\n<p class=\"wp-block-paragraph\">Root trigger:<\/p>\n<p>The CPU on the accomplice cluster was configured with C-states restricted to C1 in BIOS. It is a frequent \u201clow-latency\u201d default that&#8217;s actively flawed for AI coaching workloads. With idle cores held in C1, they continued to attract bundle energy; the busy cores feeding the GPU with kernels couldn\u2019t declare sufficient of the bundle energy funds to hit turbo. Permitting the idle cores to drop to C6 freed energy headroom, enabling the busy cores to climb to three.8 GHz and recuperate roughly 4% on this workload.<\/p>\n<p>The hypervisor housekeeping threads had been pinned to the identical bodily cores because the coaching course of\u2019s information loader staff. Contained in the VM this appeared like sporadic 50\u2013100 ms stalls within the python threads, which then propagated because the lengthy tail in step time. The repair was a cpuset separation: hypervisor and host providers on cores 0\u20137 and 56\u201363, coaching processes on the rest.<\/p>\n<p class=\"wp-block-paragraph\">End result: The 12% hole shrank to three%, with the residual traced to a special NCCL tuning subject lined within the subsequent case examine.<\/p>\n<p class=\"wp-block-paragraph\">The sample right here is that no single repair recovered the entire hole. The C-state change was the biggest single contributor at ~4%, and the remainder got here from course of isolation by way of NUMA binding. With CPU and virtualization addressed, the following ceiling is the community.<\/p>\n<h2 id=\"case_study_3_gb300_nvl72_with_nvidia_connectx-8_supernic_under-utilizing_16_tbps_fabric\" class=\"wp-block-heading\">Case examine 3: GB300 NVL72 with NVIDIA ConnectX-8 SuperNIC under-utilizing 1.6 Tbps material<\/h2>\n<p class=\"wp-block-paragraph\">Focus: ConnectX-8 SuperNIC collective tuning<\/p>\n<p class=\"wp-block-paragraph\">A GB300 NVL72 deployment with NVIDIA ConnectX-8 SuperNICs (1.6 Tbps per node) confirmed a 31% coaching efficiency hole on Nemotron-4 15B pre-training. Single-node throughput appeared wholesome; the hole appeared at 512 GPUs, the place the profiler confirmed uncovered AllGather and ReduceScatter time. That pointed to the collective path on the ConnectX-8 material slightly than compute.<\/p>\n<p class=\"wp-block-paragraph\">The investigation examined a number of variables with NCCL Exams (nccl-tests), together with iteration rely, UCX\/UCC habits, NUMA mapping, NVLS, and NCCL variations. For the workload\u2019s networking efficiency, the related tuning change was narrower: rising NCCL_IB_QPS_PER_CONNECTION to 4 from the default worth of 1.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a78f80b13cda&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a78f80b13cda\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1605\" height=\"571\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11.webp\" alt=\"Nsight Systems timelines comparing unoptimized (1.09 s) vs. optimized (0.76 s) iterations; red boxes highlight long AllGather\/ReduceScatter regions in the top trace, green boxes show shorter collectives with better overlap at QPS=4 in the bottom.\" class=\"wp-image-120325\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11.webp 1605w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-179x64.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-300x107.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-768x273.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-625x222.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-1536x546.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-645x229.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-500x178.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-160x57.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-362x129.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-309x110.png 309w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-1024x364.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-960x342.png 960w\" sizes=\"(max-width: 1605px) 100vw, 1605px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1605\" height=\"571\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11.webp\" alt=\"Nsight Systems timelines comparing unoptimized (1.09 s) vs. optimized (0.76 s) iterations; red boxes highlight long AllGather\/ReduceScatter regions in the top trace, green boxes show shorter collectives with better overlap at QPS=4 in the bottom.\" class=\"lazyload wp-image-120325\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11.webp 1605w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-179x64.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-300x107.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-768x273.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-625x222.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-1536x546.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-645x229.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-500x178.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-160x57.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-362x129.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-309x110.png 309w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-1024x364.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image2-11-960x342.png 960w\" data-sizes=\"(max-width: 1605px) 100vw, 1605px\"\/><figcaption class=\"wp-element-caption\">Determine 2. Nemotron-4 15B efficiency with NCCL QPS optimization at 512-GPU scale<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Nsight Programs hint exhibiting communication overhead uncovered at decrease QPS values, contributing to longer coaching iteration time<\/p>\n<p class=\"wp-block-paragraph\">Sign was seen in each the workload and the nccl-tests collective measurements. On the NVIDIA reference cluster, the default configuration ran at about 1.09s per iteration. With QPS=4, the identical reference workload improved to about 0.83s. Within the profile, AllGather time dropped from about 375ms to 262ms, and ReduceScatter dropped from about 389ms to 273ms. The comparability run was about 0.76s and used a special NCCL model. The remaining distinction was subsequently partly attributable to an NCCL model mismatch between the comparability and reference environments; aligning the variations narrowed the residual hole additional. As a result of NCCL model adjustments are outdoors the conventional Exemplar tuning scope, the beneficial tuning retains the deployed NCCL model unchanged.<\/p>\n<p class=\"wp-block-paragraph\">Lesson: Don\u2019t enhance QPS in every single place. QPS is fabric- and workload-dependent. On this GB300 ConnectX-8 workload, QPS=4 improved large-message AllGather and ReduceScatter habits. On different materials or message-size profiles, the identical setting could add CPU overhead with out bettering coaching throughput. The appropriate method is to check the collective on the workload\u2019s actual message sizes, sweep the setting on the goal material, and confirm the outcome within the coaching workload.<\/p>\n<h2 id=\"case_study_4_the_environment_variable_that_never_made_it_inside\" class=\"wp-block-heading\">Case examine 4: The atmosphere variable that by no means made it inside<\/h2>\n<p class=\"wp-block-paragraph\">In a virtualized B200 deployment, coaching throughput was 13%\u201353% under the NVIDIA reference although nccl-tests run on the host confirmed anticipated efficiency. Contained in the enroot workload container, AllGather and ReduceScatter had been 2\u20134\u00d7 slower, shifting the investigation from material well being to a direct comparability of the NCCL topology configuration seen on the VM and contained in the coaching job.\u00a0<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nHost (VM)                              Container (enroot)<br \/>\n\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500                         \u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500<br \/>\nNCCL_TOPO_FILE=\/and many others\/nccl\/topo.xml  \u2192NCCL_TOPO_FILE (not propagated)<br \/>\n\/and many others\/nccl\/topo.xml current       \u2192\/and many others\/nccl\/topo.xml(not mounted)<br \/>\n                                                \u2193<br \/>\n                                  NCCL falls again to auto-detection<br \/>\n                                    \u2192 13\u201353% under reference\n<\/div>\n<figure class=\"wp-block-table\">PlatformB200, virtualized stackSymptom13\u201353% under reference; AllGather\/ReduceScatter 2\u20134x slower; NCCL checks on host cross fineRoot causeNCCL_TOPO_FILE set on the VM however neither the variable nor the topology file was mounted into the enroot containerFix&#8211;mount sort=bind,supply=\/and many others\/nccl\/topo.xml,goal=\/and many others\/nccl\/topo.xml<figcaption class=\"wp-element-caption\">Desk 1. Prognosis and remediation of lacking NCCL topology configuration inside a virtualized B200 workload container<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">Lesson: Run checks from inside the identical container, launcher, and Slurm allocation that can run the benchmark not from the host. Operating echo $NCCL_TOPO_FILE &amp;&amp; cat $NCCL_TOPO_FILE contained in the job container is the quickest sanity examine. If the trail doesn\u2019t resolve, NCCL fails silently with no error making this one of many more durable gaps to diagnose with out realizing the place to look.<\/p>\n<p class=\"wp-block-paragraph\">Abstract of fixes<\/p>\n<figure class=\"wp-block-table\">CasePlatformLayerDiagnostic signalFixRecovered1GB200 NVL72 (VM)SMMUarm_smmu_cmdq_issue_cmdlist dominant in perf; multi-fold dTLB miss increaseEnable VMDQV~12percent2H100 (VM)CPU + NUMACores caught at 3.0 GHz; bimodal step time; 18% NUMA-remoteC-state tuning, cpuset isolation, numactl binding, SMT\/mitigations off9% (12-3)3GB300 NVL72NCCL concurrencyAllGather busbw at ~28 GB\/s vs ~61 GB\/s with QPS=4NCCL_IB_QPS_PER_CONNECTION replace from 1 to 4 for CX831% iter time4B200 (VM)Runtime-visible topologyHost NCCL topology appeared appropriate, however contained in the enroot container NCCL_TOPO_FILE was not propagated and \/and many others\/nccl\/topo.xml was not mounted; AllGather\/ReduceScatter had been 2-4x slowerBind-mount the topology file into the container and confirm NCCL_TOPO_FILE from contained in the job containerClosed 13-53% reference gaps<figcaption class=\"wp-element-caption\">Desk 2. Diagnostic indicators, corrective actions, and recovered efficiency throughout the 4 Exemplar Cloud case research<\/figcaption><\/figure>\n<h2 id=\"preflight_checks_before_full-scale_training_debug\" class=\"wp-block-heading\">Preflight checks earlier than full-scale coaching debug<\/h2>\n<p class=\"wp-block-paragraph\">When a cluster underperforms relative to its NVIDIA reference structure specs, these checks assist rule out frequent platform points earlier than full-scale workload tuning.<\/p>\n<figure class=\"wp-block-table\">AreaWhat to checkUseful toolsGPU and {hardware} healthClock, energy, thermal, and NVLink bandwidth consistency underneath sustained loadnvidia-smi, DCGM, dcgm-exporterGrace and VM readinessCMDQV assist, visitor web page measurement, IOMMU passthrough habits, and large-page availabilityperf, dmesg, kernel config, boot parametersCPU energy and placementBusy-core turbo, cpuset isolation, and NUMA \/ PCT binding close to GPUsturbostat, lscpu, numactl, nvidia-smi topo -mRuntime topologyTopology information, NCCL atmosphere variables, and HCA visibility contained in the job containerenv, cat $NCCL_TOPO_FILE, NCCL_DEBUG=INFOFabric collectivesAllGather and ReduceScatter habits at workload message sizesnccl-tests, workload tracesWorkload tuningPipeline parallelism, microbatch sizing, and communication overlap \u2014 solely after platform points are dominated outNsight Programs, workload logs<figcaption class=\"wp-element-caption\">Desk 3. Really useful preflight checks and diagnostic instruments for evaluating GPU well being, VM readiness, CPU placement, runtime topology, material collectives, and workload configuration earlier than full-scale coaching debugging<\/figcaption><\/figure>\n<h2 id=\"debug_early_debug_less\u00a0\" class=\"wp-block-heading\">Debug early, debug much less\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Efficiency gaps between a cloud coaching deployment and the corresponding NVIDIA reference structure are sometimes cumulative with a number of % from CPU energy settings, one other from NUMA or PCT binding, extra from a lacking kernel functionality, container-visible topology, or material configuration. These points are price checking earlier than validation as a result of they&#8217;ll flip into costly full-scale debug classes.<\/p>\n<p class=\"wp-block-paragraph\">On the similar time, preflight diagnostics don\u2019t assure an Exemplar Cloud cross. Some points solely seem within the validation workloads themselves, underneath the precise mannequin, precision, topology, container, launcher, and community situations used for the run. The sensible aim is to take away recognized platform dangers early, then use the coaching workload traces to debug the gaps that solely seem at scale.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Two AI computing clusters constructed from equivalent NVIDIA H100, GB200 NVL72, or GB300 NVL72 methods can ship materially completely different coaching throughput. We routinely see 8% to 12% gaps between accomplice deployments and the corresponding NVIDIA reference structure (RA) on the identical workload, similar mannequin, similar world batch measurement. The trigger is commonly a stack [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3523,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image6-6.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[1271,3886,3887,439,408,81,750,956],"class_list":["post-3521","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-cloud","tag-exemplar","tag-full","tag-infrastructure","tag-lessons","tag-nvidia","tag-performance","tag-unlocking"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure - Future News 24<\/title>\n<meta name=\"description\" content=\"Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-30T16:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-09T21:59:06+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image6-6.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image6-6.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure\",\"datePublished\":\"2026-07-30T16:00:00+00:00\",\"dateModified\":\"2026-08-09T21:59:06+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/\"},\"wordCount\":2289,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/image6-6.webp\",\"keywords\":[\"Cloud\",\"Exemplar\",\"Full\",\"Infrastructure\",\"lessons\",\"NVIDIA\",\"performance\",\"Unlocking\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/\",\"name\":\"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/image6-6.webp\",\"datePublished\":\"2026-07-30T16:00:00+00:00\",\"dateModified\":\"2026-08-09T21:59:06+00:00\",\"description\":\"Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/image6-6.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/image6-6.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/30\\\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure - Future News 24","description":"Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/","og_locale":"en_US","og_type":"article","og_title":"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure - Future News 24","og_description":"Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput.","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/","og_site_name":"Future News 24","article_published_time":"2026-07-30T16:00:00+00:00","article_modified_time":"2026-08-09T21:59:06+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image6-6.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image6-6.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure","datePublished":"2026-07-30T16:00:00+00:00","dateModified":"2026-08-09T21:59:06+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/"},"wordCount":2289,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image6-6.webp","keywords":["Cloud","Exemplar","Full","Infrastructure","lessons","NVIDIA","performance","Unlocking"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/","name":"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image6-6.webp","datePublished":"2026-07-30T16:00:00+00:00","dateModified":"2026-08-09T21:59:06+00:00","description":"Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput.","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image6-6.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image6-6.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/30\/nvidia-exemplar-cloud-lessons-for-unlocking-full-performance-on-ai-infrastructure\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"NVIDIA Exemplar Cloud: Classes for Unlocking Full Efficiency on AI Infrastructure"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3521","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=3521"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3521\/revisions"}],"predecessor-version":[{"id":3522,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3521\/revisions\/3522"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/3523"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=3521"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=3521"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=3521"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}