{"id":1616,"date":"2026-06-26T16:00:00","date_gmt":"2026-06-26T16:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/"},"modified":"2026-06-28T22:01:23","modified_gmt":"2026-06-28T22:01:23","slug":"creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/","title":{"rendered":"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">As context home windows develop longer, shifting massive mannequin weights effectively turns into crucial to efficiency. A typical approach to tackle that is quantization, an optimization approach that compresses mannequin weights right into a smaller knowledge format. One quantization format is NVFP4, an progressive 4-bit floating level launched with NVIDIA Blackwell structure.<\/p>\n<p class=\"wp-block-paragraph\">That\u2019s the strategy behind our new Nemotron 3 Extremely NVFP4 checkpoint: we quantized the mannequin into NVFP4 utilizing NVIDIA Mannequin Optimizer. The result&#8217;s a mannequin that achieves as much as 5.9x larger inference throughput than GLM-5.1 754B FP4 mannequin on decode-heavy workloads whereas matching BF16 accuracy throughout almost each benchmark, as proven in Determine 1.<\/p>\n<p class=\"wp-block-paragraph\">Whereas the efficiency advantages of NVFP4 are nicely understood, the method of manufacturing a high-quality NVFP4 checkpoint isn&#8217;t. This submit walks by means of how we quantized Nemotron 3 Extremely (550B) to NVFP4 with NVIDIA Mannequin Optimizer, and exhibits builders learn how to generate the most effective quantized checkpoints for their very own fashions.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4199b1a4e69&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4199b1a4e69\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"1180\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance.webp\" alt=\"Figure 1 Table mapping each Nemotron 3 Ultra layer to its BF16 baseline and quantized precision, showing mixed NVFP4, FP8, BF16, and FP16 formats by layer type.\" class=\"wp-image-119080\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-179x106.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-300x177.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-768x453.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-625x369.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-1536x907.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-645x381.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-500x295.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-152x90.png 152w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-362x214.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-186x110.png 186w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-1024x604.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-915x540.png 915w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"1180\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance.webp\" alt=\"Figure 1 Table mapping each Nemotron 3 Ultra layer to its BF16 baseline and quantized precision, showing mixed NVFP4, FP8, BF16, and FP16 formats by layer type.\" class=\"lazyload wp-image-119080\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-179x106.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-300x177.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-768x453.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-625x369.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-1536x907.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-645x381.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-500x295.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-152x90.png 152w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-362x214.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-186x110.png 186w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-1024x604.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Ultra-3-Performance-915x540.png 915w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 1. Efficiency of Nemotron 3 Extremely NVFP4 in comparison with different NVFP4 fashions<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"the_nemotron_3_ultra_nvfp4_checkpoint\" class=\"wp-block-heading\">The Nemotron 3 Extremely NVFP4 checkpoint<\/h2>\n<p class=\"wp-block-paragraph\">A typical false impression is that each layer of an NVFP4 checkpoint is saved in NVFP4. As Desk 1 exhibits, this isn\u2019t the case: completely different layers are quantized to completely different precision codecs, chosen in accordance with every layer\u2019s sensitivity to the structure and its influence on mannequin accuracy. After NVFP4 quantization, the Nemotron 3 Extremely mannequin shrinks from 1,121 GB in BF16 right down to 352.3 GB, a 3.2x discount. The payoff is substantial, chopping the {hardware} footprint in half.\u00a0<\/p>\n<figure class=\"wp-block-table\">Layer\/operatorBF16 baselineQuantized checkpoint precisionEmbedding, Output classification layer, MTP layersBF16BF16MoE routed expertsBF16NVFP4MoE shared expertsBF16FP8 per-tensorMamba mixer linearsBF16FP8 per-tensorAttention linearsBF16BF16Latent MoEBF16BF16Mamba conv1dBF16BF16KV cacheBF16FP8Mamba SSM cacheFP32FP16 with stochastic rounding<figcaption class=\"wp-element-caption\">Desk 1. BF16 baseline in comparison with the quantized checkpoint precision for every layer\/operator from the\u00a0Nemotron 3 Extremely paper<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">A key innovation of the Nemotron 3 Extremely NVFP4 is {that a} single checkpoint can run on each NVIDIA Hopper and Blackwell. It achieves this by changing the load format to match the {hardware} it runs on. On Hopper, which lacks native FP4 tensor cores, the serving framework routinely switches to W4A16. On Blackwell, it makes use of native W4A4.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Whereas W8A8 (8-bit weights, 8-bit activations) looks like the plain Hopper alternative, its bigger reminiscence footprint leaves too little headroom to suit Multi-Token Prediction (MTP). We discovered MTP may solely match alongside W4A16 (4-bit weights, 16-bit activations) so W4A16 matches or beats it throughout the board. Learn the total  Nemotron 3 Extremely technical report (Part 4.6) to be taught extra.<\/p>\n<h2 id=\"how_we_found_the_optimal_nvfp4_checkpoint\" class=\"wp-block-heading\">How we discovered the optimum NVFP4 checkpoint<\/h2>\n<p class=\"wp-block-paragraph\">Discovering an optimum NVFP4 checkpoint requires some iterations. We dive into the developer story of how we received an NVFP4 checkpoint on this part.<\/p>\n<h3 id=\"the_challenge_of_quantizing_at_fp4\" class=\"wp-block-heading\">The problem of quantizing at FP4<\/h3>\n<p class=\"wp-block-paragraph\">With FP4 quantization, there are solely 8 optimistic values [0, 0.5, 1, 1.5, 2, 3, 4, and 6] to signify a complete block of weights. We have to decide learn how to map the unique vary of values. That is managed by a scale, primarily a multiplier that determines the granularity of the illustration. Selecting a poor scale means we both waste precision on small values or clip massive values, each of which harm mannequin high quality. So how ought to we select the optimum scale issue? There are a number of approaches.<\/p>\n<h4 class=\"wp-block-heading\">Max scaling<\/h4>\n<p class=\"wp-block-paragraph\">Right here, we set the size so the biggest worth within the block maps to the utmost representable FP4 worth. Nevertheless, with the presence of a single massive weight outlier, the max scaling compresses each different worth within the block right into a slim vary, which may find yourself flushing these values to zero. This data loss might adversely have an effect on accuracy.\u00a0 Max scaling preserves the best magnitude worth within the block, with a possible aspect impact of flushing different values to zero.<\/p>\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4199b1a65ae&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4199b1a65ae\" class=\"wp-block-image size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"1290\" height=\"1058\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195.webp\" alt=\"Diagram showing six small FP32 weights collapsing to zero under absmax FP4 quantization while a single outlier (12.8) survives.\" class=\"wp-image-119094\" style=\"aspect-ratio:1.4992470299514753;width:817px;height:auto\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195.webp 1290w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-140x115.webp 140w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-300x246.webp 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-768x630.webp 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-625x513.webp 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-645x529.webp 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-366x300.webp 366w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-110x90.webp 110w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-362x297.webp 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-134x110.webp 134w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-1024x840.webp 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-658x540.webp 658w\" sizes=\"(max-width: 1290px) 100vw, 1290px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1290\" height=\"1058\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195.webp\" alt=\"Diagram showing six small FP32 weights collapsing to zero under absmax FP4 quantization while a single outlier (12.8) survives.\" class=\"lazyload wp-image-119094\" style=\"aspect-ratio:1.4992470299514753;width:817px;height:auto\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195.webp 1290w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-140x115.webp 140w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-300x246.webp 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-768x630.webp 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-625x513.webp 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-645x529.webp 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-366x300.webp 366w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-110x90.webp 110w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-362x297.webp 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-134x110.webp 134w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-1024x840.webp 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Max-Scaling-e1782328670195-658x540.webp 658w\" data-sizes=\"(max-width: 1290px) 100vw, 1290px\"\/><figcaption class=\"wp-element-caption\">Determine 2. Max (absmax) scaling on a weight block with an outlier. The size set by the outlier (12.8) compresses all six small weights to close 0 after FP4 quantization and dequantization. For 12.8, the size is absmax\/6; the outlier maps to six.0 and dequantizes to precisely that quantity<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">Strive it with NVIDIA Mannequin Optimizer:\u00a0<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\n# W4A4 \u2014 weights + activations to NVFP4 (default, max scaling)<br \/>\nmannequin = mtq.quantize(mannequin, mtq.NVFP4_DEFAULT_CFG, forward_loop=forward_loop)\n<\/div>\n<p class=\"wp-block-paragraph\">Max scaling (additionally referred to as absmax, for the reason that scale is ready solely by the block\u2019s absolute most) is the best possibility, however that sensitivity to outliers makes it not often the most effective one.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">That is precisely the hole we hit on our prior mannequin, NVIDIA Nemotron 3 Tremendous: naive absmax NVFP4 PTQ left an accuracy hole, so the staff evaluated a variety of other calibration methods that don\u2019t let a single outlier dictate the size, from imply squared error (MSE)-based weight scaling to GPTQ, an environment friendly methodology that makes use of second-order data to encode weights.<\/p>\n<figure class=\"wp-block-table\">AlgorithmDetailsMMLU-ProGPQALiveCodeBenchAA-LCRBF16\u201483.4979.9272.90753.00Default NVFP4 PTQ (Baseline algorithm)Static per-tensor scales are computed utilizing max-value calibration; per-block scales are computed dynamically from block most values.82.9979.2970.1855.50Weight per-block scales minimizing MSEWeight per-block scales are swept to attenuate per-block MSE.83.3179.9271.3756.75Weight per-block scales to attenuate output MSEWeight per-block scales are swept independently to attenuate GEMM output MSE.83.0578.9871.0057.06GPTQGPTQ (Frantar et al., 2023) is used for weight quantization.83.1180.0569.7957.87<figcaption class=\"wp-element-caption\">Desk 2. Experiment outcomes for Nemotron 3 Tremendous quantization. The staff tried a number of quantization strategies and evaluated the accuracy change throughout 4 duties. For extra data, see the Nemotron 3 Tremendous paper<\/figcaption><\/figure>\n<h3 id=\"mean_squared_error_scaling\" class=\"wp-block-heading\">Imply squared error scaling<\/h3>\n<p class=\"wp-block-paragraph\">One other strategy is imply squared error (MSE) scaling, which searches for the size that minimizes common reconstruction error throughout the entire block.<\/p>\n<p class=\"wp-block-paragraph\">Nevertheless, decrease MSE doesn&#8217;t all the time translate to higher mannequin accuracy. MSE calibration lowered per-tensor weight error by 27.1% over four-over-six scaling in our Nemotron 3 Extremely experiments, but produced no constant enchancment on downstream benchmarks.<\/p>\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4199b1a7861&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4199b1a7861\" class=\"wp-block-image size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1916\" height=\"1210\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760.webp\" alt=\"Diagram showing MSE-scaled FP4 quantization preserving small weights while clipping the outlier from 12.8 down to 2.0.\" class=\"wp-image-119083\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760.webp 1916w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-179x113.webp 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-300x189.webp 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-768x485.webp 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-625x395.webp 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-1536x970.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-645x407.webp 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-475x300.webp 475w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-143x90.webp 143w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-362x229.webp 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-174x110.webp 174w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-1024x647.webp 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-855x540.webp 855w\" sizes=\"(max-width: 1916px) 100vw, 1916px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1916\" height=\"1210\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760.webp\" alt=\"Diagram showing MSE-scaled FP4 quantization preserving small weights while clipping the outlier from 12.8 down to 2.0.\" class=\"lazyload wp-image-119083\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760.webp 1916w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-179x113.webp 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-300x189.webp 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-768x485.webp 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-625x395.webp 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-1536x970.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-645x407.webp 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-475x300.webp 475w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-143x90.webp 143w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-362x229.webp 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-174x110.webp 174w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-1024x647.webp 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/MSE-Scaling-1-e1782328711760-855x540.webp 855w\" data-sizes=\"(max-width: 1916px) 100vw, 1916px\"\/><figcaption class=\"wp-element-caption\">Determine 3. Illustration of MSE-based scaling on the identical block preserves the majority of small weights with usable decision whereas saturating the outlier right down to 2.0<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">Strive MSE-based scaling\u00a0with NVIDIA Mannequin Optimizer:\u00a0<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nmannequin = mtq.quantize(mannequin, mtq.NVFP4_W4A4_WEIGHT_MSE_FP8_SWEEP_CFG, forward_loop=forward_loop)\n<\/div>\n<p class=\"wp-block-paragraph\">For our earlier mannequin, NVIDIA Nemotron 3 Tremendous, the ultimate quantization recipe mixed MSE-based block scaling for weights with a per-tensor FP8 sweep and dynamic max-based scaling for activations. Combining MSE weights with the FP8 activation sweep gave the most effective accuracy-to-size tradeoff of all the things we tried, and it turned our optimum NVFP4 configuration for Tremendous.<\/p>\n<p class=\"wp-block-paragraph\">Max and MSE scaling each choose a scale to attenuate total rounding error, however neither pays consideration to the place the error comes from on the grid. For Nemotron 3 Extremely, we used a scaling methodology that chooses the vary primarily based on the error from the hole within the grid.<\/p>\n<h3 id=\"four-over-six_scaling\" class=\"wp-block-heading\">4-over-six scaling<\/h3>\n<p class=\"wp-block-paragraph\">Keep in mind how NVFP4 can solely signify 8 optimistic values: 0, 0.5, 1, 1.5, 2, 3, 4, and 6. Discover that after 4, the subsequent worth jumps straight to six. Any weight that falls in that vary will get rounded aggressively to both 4 or 6, generally incurring over 13% error on a single worth.<\/p>\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4199b1a8636&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4199b1a8636\" class=\"wp-block-image size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1409\" height=\"858\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding.webp\" alt=\"Range showing the gap of FP4 representable values between 4-6. Bottom : FP4 number line showing the 4-to-6 gap alongside a perplexity curve that spikes as the quantization threshold nears maximal values.&#10;\" class=\"wp-image-119084\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding.webp 1409w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-179x109.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-300x183.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-768x468.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-625x381.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-645x393.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-493x300.png 493w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-148x90.png 148w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-362x220.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-181x110.png 181w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-1024x624.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-887x540.png 887w\" sizes=\"(max-width: 1409px) 100vw, 1409px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1409\" height=\"858\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding.webp\" alt=\"Range showing the gap of FP4 representable values between 4-6. Bottom : FP4 number line showing the 4-to-6 gap alongside a perplexity curve that spikes as the quantization threshold nears maximal values.&#10;\" class=\"lazyload wp-image-119084\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding.webp 1409w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-179x109.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-300x183.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-768x468.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-625x381.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-645x393.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-493x300.png 493w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-148x90.png 148w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-362x220.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-181x110.png 181w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-1024x624.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Error-Rounding-887x540.png 887w\" data-sizes=\"(max-width: 1409px) 100vw, 1409px\"\/><figcaption class=\"wp-element-caption\">Determine 4. NVFP4 error comes from rounding near-maximal values. High: FP4 representable values go away a spot between 4 and 6, and rounding into that hole drives a lot of the NVFP4 error. Backside: from Prepare dinner et al. (2026), simulated quantization on Llama-3.1-8B FP4 exhibits how the quantization threshold impacts error<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">4-over-six fixes this as every block of weights independently chooses between scaling to a most of M=4 or M=6, choosing whichever minimizes reconstruction error. 4-over-six works on weights and falls again to the default NVFP4 on activations.<\/p>\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4199b1a92cc&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4199b1a92cc\" class=\"wp-block-image size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1596\" height=\"592\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling.webp\" alt=\" Two example blocks comparing M=6 and M=4 FP4 scaling, with mean squared error showing each block prefers a different grid.&#10;\" class=\"wp-image-119085\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling.webp 1596w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-179x66.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-300x111.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-768x285.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-625x232.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-1536x570.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-645x239.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-500x185.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-160x59.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-362x134.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-297x110.png 297w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-1024x380.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-960x356.png 960w\" sizes=\"(max-width: 1596px) 100vw, 1596px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1596\" height=\"592\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling.webp\" alt=\" Two example blocks comparing M=6 and M=4 FP4 scaling, with mean squared error showing each block prefers a different grid.&#10;\" class=\"lazyload wp-image-119085\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling.webp 1596w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-179x66.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-300x111.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-768x285.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-625x232.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-1536x570.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-645x239.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-500x185.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-160x59.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-362x134.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-297x110.png 297w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-1024x380.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Block-Scaling-960x356.png 960w\" data-sizes=\"(max-width: 1596px) 100vw, 1596px\"\/><figcaption class=\"wp-element-caption\">Determine 5. M=4 vs. M=6 block scaling (four-over-six). Two pattern weight blocks are quantized at M=6 and M=4 with their FP4 codes and imply squared error, exhibiting one block favors M=6, and the opposite favors M=4<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">When M=6 wins with block: [2, 4, 5.9, 6]\u00a0<\/p>\n<p class=\"wp-block-paragraph\">At a scale of M=1, values 2, 4, and 6 map precisely onto FP4 grid factors and solely 5.9 rounds to six at negligible value. Scaling to M=4 pushes 2 to 2.25 and 4 to 4.5, introducing error.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">When M=4 wins with block: [10, 20, 30, 40]<\/p>\n<p class=\"wp-block-paragraph\">Scaling to M=6 maps 30 to 4.62, which rounds right down to 4, a 13% error. Scaling to M=4 as a substitute maps 10, 20, 30, and 40 precisely onto 1, 2, 3, and 4, with zero rounding error throughout your entire block. MSE: 4.33 vs 0.0.<\/p>\n<p class=\"wp-block-paragraph\">4-over-six was used to set the FP4 routed-expert weight scales in Nemotron 3 Extremely, elevating the worldwide per-tensor weight scale by 1.75x, and with every microblock choosing the M=4 or M=6 grid. Throughout all 49,152 projection weights within the mannequin\u2019s 48 MoE knowledgeable layers, it minimize the median reconstruction MSE by 16.4% in comparison with normal max calibration, and delivered the most effective downstream outcome within the balanced 5.03-BPE setting: 98.5% median restoration relative to BF16, forward of max (96.8%) and MSE (98.4%).<\/p>\n<p class=\"wp-block-paragraph\">Strive four-over-six with NVIDIA Mannequin Optimizer:\u00a0<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nmannequin = mtq.quantize(mannequin,<br \/>\nmtq.NVFP4_FOUR_OVER_SIX_CFG, forward_loop=forward_loop)\n<\/div>\n<p class=\"wp-block-paragraph\">NVFP4_FOUR_OVER_SIX_CFG might be launched on the upcoming 0.46 NVIDIA Mannequin Optimizer in July. View the Nemotron 3 Extremely PTQ instance.\u00a0<\/p>\n<h3 id=\"bits-per-element\" class=\"wp-block-heading\">Bits-per-element<\/h3>\n<p class=\"wp-block-paragraph\">Efficient bits-per-element (BPE) refers back to the common variety of bits required to retailer all weights of the mannequin. A mannequin with all BF16 weights makes use of 16 efficient bits-per-element, whereas a half-FP8, half-BF16 mannequin makes use of solely 12. NVFP4 provides per-block and per-tensor scaling overhead, bringing its minimal to 4.5 efficient bits-per-element. The per-tensor scale\u2019s 32 bits are amortized throughout the total tensor and is assumed to be negligible within the total BPE calculation.<\/p>\n<p class=\"wp-block-paragraph\">The objective is to seek for the quantization configuration that pushes efficient BPE as little as potential with out sacrificing accuracy. That is tough as a result of layers should not equally strong. Some are delicate to quantization and should keep in larger precision, which raises the efficient BPE. Since every layer could be quantized at a special stage or left unquantized, the variety of potential combos grows exponentially, making an exhaustive search impractical and a wiser technique essential.<\/p>\n<p>NVIDIA Mannequin Optimizer AutoQuantize (mtq.auto_quantize) does it for you. As a substitute of a set config, you give it a goal bit finances (for instance auto_quantize_bits=4.8) and a listing of candidate codecs, similar to NVFP4_DEFAULT_CFG and FP8_DEFAULT_CFG. It then scores every layer\u2019s sensitivity and searches for the per-layer format project that meets the finances at the most effective accuracy, conserving essentially the most delicate layers within the higher-precision format or skipping them solely.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nimport modelopt.torch.quantization as mtq<\/p>\n<p>mannequin, search_state = mtq.auto_quantize(<br \/>\n    mannequin,<br \/>\n    constraints={&#8220;auto_quantize_bits&#8221;: 4.8},<br \/>\n    quantization_formats=[&#8220;NVFP4_DEFAULT_CFG&#8221;, &#8220;FP8_DEFAULT_CFG&#8221;],<br \/>\n    data_loader=calib_dataloader,<br \/>\n    forward_step=forward_step,<br \/>\n    loss_func=loss_func,\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">To search out the precise bits-per-element for Nemotron 3 Extremely, we swept over 5 working factors starting from 4.85 to 7.19 efficient bits-per-element, evaluating accuracy over a number of benchmarks in Desk 3. The important thing sign got here from AA-LCR, the place going from 4.85 to five.03 improved the benchmark by 2.4 factors, and benchmark efficiency then flattened once more past 5.03. This makes 5.03 BPE the candy spot.<\/p>\n<figure class=\"wp-block-table aligncenter\">\u00a0Quantization (bits-per-element)TaskMetric4.855.03\u20205.255.437.19CodingSciCodepass@1 (avg-16), subtask acc43.8243.8843.4543.2743.44Scientific ReasoningGPQA Diamondpass@1 (avg-32), sym. correct84.6684.3384.7584.1284.52HLEpass@1, choose correct24.2424.8425.0024.9825.44CritPtpass@1 (avg-8), accuracy3.043.935.184.824.46GeneralAA-Omnisciencepass@1 (avg-20), choose correct29.2129.7529.1829.2929.00pass@1 (avg-20), non-hallucination54.1351.5951.8451.7052.81IFBenchpass@1 (avg-8), avg. score79.3479.2679.8379.5379.83Long ContextAA-LCRpass@1 (avg-16), choose correct62.2564.6964.1964.9465.00<figcaption class=\"wp-element-caption\">Desk 3. Accuracy in comparison with efficient bits-per-element from the Nemotron 3 Extremely paper<\/figcaption><\/figure>\n<h2 id=\"how_we_quantized_nemotron_3_ultra_to_nvfp4_with_model_optimizer\u00a0\" class=\"wp-block-heading\">How we quantized Nemotron 3 Extremely to NVFP4 with Mannequin Optimizer\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Not like Nemotron 3 Tremendous 120B, Nemotron 3 Extremely is a 550B mannequin, so it advantages considerably from parallelizing the quantization course of. For that reason, we help two quantization paths.<\/p>\n<figure class=\"wp-block-table aligncenter\">Each paths are powered by NVIDIA Mannequin OptimizerMetricHugging Face TransformersMegatron-LMCompute4 \u00d7 B30016 \u00d7 B300;Knowledgeable parallelism = knowledge parallelism = 16Model loading time40 min&lt; 2 minModel loading and calibration time85 min9 minExport42 min33 minTotal time120 min45 min<figcaption class=\"wp-element-caption\">Desk 4. Quantization time evaluating Hugging Face Transformers to Megatron-LM<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">Quantizing Nemotron 3 Extremely to NVFP4 follows the NVIDIA ModelOpt post-training quantization (PTQ) pipeline in NVIDIA Megatron-LM. With the parallel route, the pretrained checkpoint is first transformed to Megatron-LM format after which quantized with a single name to quantize.sh, passing an NVFP4 quantization config because the recipe. On the backend, Megatron-LM shards the mannequin throughout GPUs with knowledgeable and knowledge parallelism (EP = DP = 16 on 16\u00d7B300s), so the calibration ahead go runs distributed throughout all units. This reduces load and calibration from ~85 minutes to ~9.<\/p>\n<p class=\"wp-block-paragraph\">Calibration runs nemotron-post-training-dataset-v2 to suit the per-block scales, and the precision coverage is solely config-driven. Choose it by passing a config to quantize.sh. Both a built-in identify (e.g., NVFP4_DEFAULT_CFG, FP8_DEFAULT_CFG) or a YAML recipe path, which is what in the end will get handed to mtq.quantize(mannequin, config, forward_loop) to put in the quantizers and run calibration.<\/p>\n<p class=\"wp-block-paragraph\">Strive 4-Over-Six Scaling with NVIDIA Mannequin Optimizer:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nHF_MODEL_CKPT=nvidia\/NVIDIA-Nemotron-3-Extremely-550B-A55B-BF16<\/p>\n<p># Step 1 \u2014 Quantize to NVFP4<br \/>\nTP=4<br \/>\nMLM_MODEL_SAVE=\/tmp\/Nemotron-3-Ultra_quant<br \/>\n.\/quantize.sh nvidia\/NVIDIA-Nemotron-3-Extremely-550B-A55B-BF16 huggingface\/fashions\/nvidia\/Nemotron-3-Extremely-550B-A55B\/ptq\/nvfp4-4o6<\/p>\n<p># Step 2 \u2014 Export the quantized checkpoint<br \/>\nPP=1<br \/>\nMLM_MODEL_CKPT=\/tmp\/Nemotron-3-Ultra_quant<br \/>\nEXPORT_DIR=\/tmp\/Nemotron-3-Ultra_NVFP4_46_HF<br \/>\n.\/export.sh nvidia\/NVIDIA-Nemotron-3-Extremely-550B-A55B-BF16<\/p>\n<\/div>\n<p class=\"wp-block-paragraph\">NVFP4_FOUR_OVER_SIX_CFG help for four-over-six is touchdown in NVIDIA Mannequin Optimizer 0.46. The Nemotron 3 Extremely recipe for four-over-six is offered on GitHub. 4-over-six works on weights and falls again to default NVFP4 on activations.<\/p>\n<h2 id=\"customizing_quantization_configs\" class=\"wp-block-heading\">Customizing quantization configs<\/h2>\n<p class=\"wp-block-paragraph\">NVIDIA Mannequin Optimizer is constructed to be customizable with completely different quantization configs. The built-in NVFP4 configs vary from NVFP4_DEFAULT_CFG, which quantizes broadly, to extra selective presets like NVFP4_MLP_ONLY_CFG, NVFP4_EXPERTS_ONLY_CFG, and NVFP4_OMLP_ONLY_CFG that limit FP4 to the MLP and knowledgeable layers whereas conserving the delicate consideration projections in larger precision.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Below the hood, a config is an ordered record of guidelines matched towards module-name patterns, and mtq.quantize() applies them. Weight quantization is ruled by guidelines concentrating on the *weight_quantizer sample, the place you set the format (for NVFP4, E2M1 components with 16-wide blocks and E4M3 block scales), whereas activation quantization is ruled by separate guidelines on the *input_quantizer sample.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Because the two are impartial, you&#8217;ll be able to quantize weights solely or weights and activations collectively, and you may carve out exceptions for particular modules by appending guidelines that disable them. For something past the built-in presets, you&#8217;ll be able to write a full YAML recipe and cargo it with &#8211;recipe, which then absolutely defines the quant config.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The next Nemotron-3 Extremely recipe applies NVFP4 with four-over-six to the routed-expert weights, retains the shared specialists and Mamba projections in FP8, makes use of an FP8 KV cache, and leaves all the things else in BF16. The whole recipe ships with NVIDIA Mannequin Optimizer\u2019s recipe library: nvfp4-4o6.yaml<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\n# Nemotron 3 Extremely NVFP4 mixed-precision recipe with 4-Over-Six (4\/6)<br \/>\n# Instance recipe for HuggingFace fashions, for Megatron-compatible recipe see the total recipe hyperlink<\/p>\n<p>quantize:<br \/>\n  algorithm:<br \/>\n    methodology: mse<br \/>\n    fp8_scale_sweep: false<br \/>\n    start_multiplier: 1.0   # M=6 (hold amax)<br \/>\n    stop_multiplier: 1.5    # M=4 (amax x 6\/4)<br \/>\n    step_size: 0.5          # candidates [1.0, 1.5]<\/p>\n<p>  quant_cfg:<br \/>\n    # Disable all the things by default; later guidelines re-enable particular modules.<br \/>\n    &#8211; quantizer_name: &#8216;*&#8217;<br \/>\n      allow: false<\/p>\n<p>    # MoE routed specialists -&gt; NVFP4 W4A4, block 16, e4m3 block scale.<br \/>\n    # 4\/6 adaptive block scaling on weights solely; not actvivations<br \/>\n    # HF names: spine.layers.*.mixer.specialists.*.{up,down}_proj<br \/>\n    &#8211; quantizer_name: &#8216;*mixer.specialists.*weight_quantizer&#8217;<br \/>\n      allow: true<br \/>\n      cfg:<br \/>\n        block_sizes: {-1: 16, sort: static, scale_bits: e4m3, four_over_six: true}<br \/>\n        num_bits: e2m1<br \/>\n    &#8211; quantizer_name: &#8216;*mixer.specialists.*input_quantizer&#8217;<br \/>\n      allow: true<br \/>\n      cfg:<br \/>\n        block_sizes: {-1: 16, sort: dynamic, scale_bits: e4m3}<br \/>\n        num_bits: e2m1<\/p>\n<p>    # Shared specialists + Mamba in\/out_proj -&gt; FP8 per-tensor (weights+activations).<br \/>\n    &#8211; quantizer_name: &#8216;*mixer.shared_experts*&#8217;<br \/>\n      allow: true<br \/>\n      cfg: {num_bits: e4m3, axis: null}<br \/>\n    &#8211; quantizer_name: &#8216;*mixer.in_proj*&#8217;<br \/>\n      allow: true<br \/>\n      cfg: {num_bits: e4m3, axis: null}<br \/>\n    &#8211; quantizer_name: &#8216;*mixer.out_proj*&#8217;<br \/>\n      allow: true<br \/>\n      cfg: {num_bits: e4m3, axis: null}<\/p>\n<p>    # KV cache -&gt; FP8.<br \/>\n    &#8211; quantizer_name: &#8216;*[kv]_bmm_quantizer&#8217;<br \/>\n      allow: true<br \/>\n      cfg: {num_bits: e4m3}\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">Whereas we walked by means of this on Nemotron 3 Extremely, the identical pipeline works with any Hugging Face mannequin checkpoint. Merely level Mannequin Optimizer at a mannequin card from the Hub or a neighborhood path, choose a config (a built-in preset or your personal recipe), and run the identical quantize and export steps.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nimport modelopt.torch.quantization as mtq<br \/>\nfrom modelopt.torch.export import export_hf_checkpoint<br \/>\nfrom transformers import AutoModelForCausalLM<\/p>\n<p>mannequin = AutoModelForCausalLM.from_pretrained(&#8220;&#8221;)<\/p>\n<p># Calibrate + quantize with the config of your alternative<br \/>\nmannequin = mtq.quantize(mannequin, mtq.NVFP4_DEFAULT_CFG, forward_loop)<\/p>\n<p># Export a unified HF checkpoint for TRT-LLM \/ vLLM \/ SGLang<br \/>\nexport_hf_checkpoint(mannequin, export_dir=&#8221;&#8221;)\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">Strive the one-click launcher<\/p>\n<p class=\"wp-block-paragraph\">To simplify deployment, the Mannequin Optimizer launcher automates your entire Extremely PTQ and export workflow. After finishing the setup steps within the launcher README, the workflow could be launched by way of the Nemotron 3 Extremely YAML recipe with a single command from a neighborhood machine:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nuv run launch.py &#8211;yaml examples\/nvidia\/NVIDIA-Nemotron-3-Extremely-550B-A55B-BF16\/megatron_lm_ptq.yaml &#8211;yes\n<\/div>\n<p class=\"wp-block-paragraph\">As soon as launched, the workflow handles the remaining quantization and export steps routinely, assuming entry to a Slurm cluster with ample GPU sources. This instance was validated on 4 nodes, every outfitted with 4 NVIDIA Blackwell GPUs.<\/p>\n<p class=\"wp-block-paragraph\">For smaller-scale deployments, a PTQ instance can also be obtainable for Nemotron-3 Tremendous:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nuv run launch.py &#8211;yaml examples\/nvidia\/NVIDIA-Nemotron-3-Tremendous-120B-A12B-BF16\n<\/div>\n<h2 id=\"get_started\" class=\"wp-block-heading\">Get began<\/h2>\n<p class=\"wp-block-paragraph\">This course of could be reproduced utilizing the total recipe obtainable within the open-source NVIDIA Mannequin Optimizer GitHub repository. Mannequin Optimizer is a community-driven undertaking, and contributions are inspired. Points could be filed to report bugs or request options, the undertaking roadmap could be reviewed for upcoming work, and pull requests could be submitted to contribute enhancements. See CONTRIBUTING.md for contribution pointers and getting-started data.<\/p>\n<p class=\"wp-block-paragraph\">Be taught extra with the next sources:<\/p>\n<h2 id=\"acknowledgments\" class=\"wp-block-heading\">Acknowledgments<\/h2>\n<p class=\"wp-block-paragraph\">This work wouldn&#8217;t have been potential with out the shut collaboration between the NVIDIA Mannequin Optimizer staff and the Nemotron staff. We thank the engineers throughout each groups who contributed to the quantization pipeline, analysis infrastructure, and mannequin coaching. Particular due to the Megatron-LM staff for enabling distributed quantization at scale, and to the Nemotron staff for the benchmark suite used to validate the FP4 recipes. We additionally thank the broader NVIDIA Analysis and Utilized Deep Studying groups for his or her continued help and suggestions all through this undertaking.<\/p>\n<p class=\"wp-block-paragraph\">Particularly, we thank Asma Kuriparambil Thekkumpate, Jenny Chen, and Jinhang Choi for main the implementation of the NVFP4 quantization on Nemotron 3 Extremely.\u00a0<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>As context home windows develop longer, shifting massive mannequin weights effectively turns into crucial to efficiency. A typical approach to tackle that is quantization, an optimization approach that compresses mannequin weights right into a smaller knowledge format. One quantization format is NVFP4, an progressive 4-bit floating level launched with NVIDIA Blackwell structure. That\u2019s the strategy [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1618,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Model-Optimizer.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[2077,2076,105,48,1034,81,2078,82],"class_list":["post-1616","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-checkpoint","tag-creating","tag-model","tag-nemotron","tag-nvfp4","tag-nvidia","tag-optimizer","tag-ultra"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer - Future News 24<\/title>\n<meta name=\"description\" content=\"As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization&#8230;\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer - Future News 24\" \/>\n<meta property=\"og:description\" content=\"As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization&#8230;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-26T16:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-28T22:01:23+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Model-Optimizer.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Model-Optimizer.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"15 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer\",\"datePublished\":\"2026-06-26T16:00:00+00:00\",\"dateModified\":\"2026-06-28T22:01:23+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/\"},\"wordCount\":3104,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/Model-Optimizer.webp\",\"keywords\":[\"Checkpoint\",\"Creating\",\"Model\",\"Nemotron\",\"NVFP4\",\"NVIDIA\",\"Optimizer\",\"Ultra\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/\",\"name\":\"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/Model-Optimizer.webp\",\"datePublished\":\"2026-06-26T16:00:00+00:00\",\"dateModified\":\"2026-06-28T22:01:23+00:00\",\"description\":\"As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization&#8230;\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/Model-Optimizer.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/Model-Optimizer.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/26\\\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer - Future News 24","description":"As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization&#8230;","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/","og_locale":"en_US","og_type":"article","og_title":"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer - Future News 24","og_description":"As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization&#8230;","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/","og_site_name":"Future News 24","article_published_time":"2026-06-26T16:00:00+00:00","article_modified_time":"2026-06-28T22:01:23+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Model-Optimizer.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Model-Optimizer.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"15 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer","datePublished":"2026-06-26T16:00:00+00:00","dateModified":"2026-06-28T22:01:23+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/"},"wordCount":3104,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Model-Optimizer.webp","keywords":["Checkpoint","Creating","Model","Nemotron","NVFP4","NVIDIA","Optimizer","Ultra"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/","name":"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Model-Optimizer.webp","datePublished":"2026-06-26T16:00:00+00:00","dateModified":"2026-06-28T22:01:23+00:00","description":"As context windows grow longer, moving large model weights efficiently becomes critical to performance. A common way to address this is quantization&#8230;","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Model-Optimizer.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/Model-Optimizer.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/26\/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Creating the NVIDIA Nemotron 3 Extremely NVFP4 Checkpoint with NVIDIA Mannequin Optimizer"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1616","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=1616"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1616\/revisions"}],"predecessor-version":[{"id":1617,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1616\/revisions\/1617"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/1618"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=1616"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=1616"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=1616"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}