{"id":2786,"date":"2026-07-23T00:00:00","date_gmt":"2026-07-23T00:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/"},"modified":"2026-07-24T12:59:07","modified_gmt":"2026-07-24T12:59:07","slug":"nunchaku-diffusers","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/","title":{"rendered":"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<div class=\"not-prose\">\n<div class=\"SVELTE_HYDRATER contents\" data-target=\"BlogAuthorsByline\" data-props=\"{&quot;authors&quot;:[{&quot;author&quot;:{&quot;_id&quot;:&quot;655647bf2f76548766fce60f&quot;,&quot;avatarUrl&quot;:&quot;https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/655647bf2f76548766fce60f\/mZhWuhJb-bzlzp-C3VoO_.jpeg&quot;,&quot;fullname&quot;:&quot;Pham Hong Vinh&quot;,&quot;name&quot;:&quot;rootonchair&quot;,&quot;type&quot;:&quot;user&quot;,&quot;isPro&quot;:true,&quot;isHf&quot;:false,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;followerCount&quot;:13,&quot;isUserFollowing&quot;:false},&quot;guest&quot;:true},{&quot;author&quot;:{&quot;_id&quot;:&quot;5f7fbd813e94f16a85448745&quot;,&quot;avatarUrl&quot;:&quot;https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/1649681653581-5f7fbd813e94f16a85448745.jpeg&quot;,&quot;fullname&quot;:&quot;Sayak Paul&quot;,&quot;name&quot;:&quot;sayakpaul&quot;,&quot;type&quot;:&quot;user&quot;,&quot;isPro&quot;:true,&quot;isHf&quot;:true,&quot;isHfAdmin&quot;:false,&quot;isMod&quot;:false,&quot;followerCount&quot;:1003,&quot;isUserFollowing&quot;:false}}],&quot;translators&quot;:[],&quot;proofreaders&quot;:[],&quot;lang&quot;:&quot;en&quot;}\">\n<div class=\"not-prose\">\n<div class=\"mb-12 flex flex-wrap items-center gap-x-5 gap-y-3.5\">\n<div class=\"flex items-center font-sans leading-tight\"><span class=\"inline-block \"><span class=\"contents\"><img decoding=\"async\" class=\"rounded-full! m-0 mr-2.5 size-9 sm:mr-3 sm:size-12\" alt=\"Pham Hong Vinh's avatar\" src=\"https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/655647bf2f76548766fce60f\/mZhWuhJb-bzlzp-C3VoO_.jpeg\"\/><\/span> <\/span> <\/div>\n<div class=\"flex items-center font-sans leading-tight\"><span class=\"inline-block \"><span class=\"contents\"><img decoding=\"async\" class=\"rounded-full! m-0 mr-2.5 size-9 sm:mr-3 sm:size-12\" alt=\"Sayak Paul's avatar\" src=\"https:\/\/cdn-avatars.huggingface.co\/v1\/production\/uploads\/1649681653581-5f7fbd813e94f16a85448745.jpeg\"\/><\/span> <\/span> <\/div>\n<\/div><\/div>\n<\/div>\n<\/div>\n<p>Massive diffusion transformers can create gorgeous photographs (and even movies, audio snippets, and now textual content), however loading a contemporary text-to-image mannequin in BF16 precision typically requires 20-30 GB of VRAM, which places these fashions out of attain of most client GPUs. Quantization is a robust resolution to this downside, and Diffusers already integrates a number of quantization backends reminiscent of bitsandbytes, GGUF, torchao, and Quanto, which we lined in Exploring Quantization Backends in Diffusers.<\/p>\n<p>Most of those backends are weight-only. Because of this they retailer the weights in low precision and dequantize them again to excessive precision at compute time. This reduces reminiscence utilization considerably, nevertheless it normally doesn&#8217;t make inference sooner, and might even add a small latency overhead.<\/p>\n<p>SVDQuant, the quantization methodology behind the favored Nunchaku inference engine, takes a distinct strategy. It runs the principle transformer layers with 4-bit weights and activations (W4A4), decreasing reminiscence whereas additionally dashing up the denoising loop. The small print are lined under, however till now, utilizing these checkpoints required a separate inference library.<\/p>\n<p>With present Diffusers, loading a Nunchaku checkpoint is so simple as calling from_pretrained(), with no native CUDA compilation required due to the kernels package deal. As well as, the companion diffuse-compressor toolkit helps you to quantize new architectures your self and publish them as common Diffusers repositories.<\/p>\n<figure class=\"image text-center\">\n  <img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/nunchaku-diffusers\/contact_sheet_top3_metrics_bold.png\" alt=\"Nunchaku Lite image quality and performance comparison\"\/><br \/>\n<\/figure>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tDesk of Contents<br \/>\n\t<\/span><br \/>\n<\/h2>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tGetting began with Nunchaku Lite<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>First, set up the necessities. You want a latest model of Diffusers and the Hugging Face kernels package deal:<\/p>\n<p>pip set up -U diffusers transformers speed up kernels bitsandbytes<\/p>\n<p>Then load a pre-quantized pipeline like every other Diffusers mannequin:<\/p>\n<p><span class=\"hljs-keyword\">import<\/span> torch<br \/>\n<span class=\"hljs-keyword\">from<\/span> diffusers <span class=\"hljs-keyword\">import<\/span> ErnieImagePipeline<\/p>\n<p>pipe = ErnieImagePipeline.from_pretrained(<br \/>\n    <span class=\"hljs-string\">&#8220;lite-infer\/ERNIE-Picture-Turbo-nunchaku-lite-nvfp4_r32-bnb4-text-encoder&#8221;<\/span>,<br \/>\n    torch_dtype=torch.bfloat16,<br \/>\n).to(<span class=\"hljs-string\">&#8220;cuda&#8221;<\/span>)<\/p>\n<p>picture = pipe(<br \/>\n    immediate=<span class=\"hljs-string\">&#8220;A cinematic portrait of a crimson fox in a misty forest at dawn, &#8220;<\/span><br \/>\n           <span class=\"hljs-string\">&#8220;detailed fur, volumetric mild&#8221;<\/span>,<br \/>\n    top=<span class=\"hljs-number\">1024<\/span>,<br \/>\n    width=<span class=\"hljs-number\">1024<\/span>,<br \/>\n    num_inference_steps=<span class=\"hljs-number\">8<\/span>,<br \/>\n    guidance_scale=<span class=\"hljs-number\">1.0<\/span>,<br \/>\n    generator=torch.Generator(<span class=\"hljs-string\">&#8220;cuda&#8221;<\/span>).manual_seed(<span class=\"hljs-number\">42<\/span>),<br \/>\n).photographs[<span class=\"hljs-number\">0<\/span>]<br \/>\npicture.save(<span class=\"hljs-string\">&#8220;output.png&#8221;<\/span>)<\/p>\n<figure class=\"image text-center\">\n  <img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/nunchaku-diffusers\/fox_bf16_vs_nunchaku_no_metrics.png\" alt=\"BF16 and Nunchaku Lite outputs for a red fox prompt\"\/><br \/>\n<\/figure>\n<p>No customized pipeline class or separate inference engine is required, and there may be nothing to compile domestically. The NVFP4 kernels are downloaded from the Hub via the Nunchaku Lite kernels web page the primary time they&#8217;re used. This checkpoint pairs a Nunchaku NVFP4 transformer with a bitsandbytes NF4 textual content encoder, and generates a 1024&#215;1024 picture in about 1.7 seconds on an RTX 5090 with a peak reminiscence utilization of about 12 GB, in contrast with about 24 GB for the BF16 pipeline. You will discover extra particulars concerning the Nunchaku Lite checkpoint format within the official Diffusers documentation.<\/p>\n<blockquote class=\"note\">\n<p>NVFP4 checkpoints require an NVIDIA Blackwell GPU (RTX 50 sequence, RTX PRO 6000, B200). For earlier generations, use the INT4 variants. See the {hardware} help desk under for particulars.<\/p>\n<\/blockquote>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tBackground: SVDQuant and Nunchaku<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>SVDQuant is the quantization methodology behind Nunchaku, its reference CUDA inference engine. Normal 4-bit quantization is troublesome for diffusion transformers as a result of each weights and activations comprise massive outliers. SVDQuant handles this by transferring activation outliers into the weights, representing the toughest a part of every weight matrix with a small 16-bit low-rank department, and quantizing the remaining residual to 4 bits. Nunchaku makes this quick with fused kernels for the 4-bit path and the low-rank department.<\/p>\n<figure class=\"image text-center\">\n  <img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/nunchaku-diffusers\/svdquant_kernel_fusion.png\" alt=\"Nunchaku kernel fusion: the low-rank down projection is fused with input quantization, and the low-rank up projection is fused with the 4-bit matmul\"\/><figcaption>Nunchaku fuses the low-rank down projection with the quantization kernel and the low-rank up projection with the 4-bit compute kernel, eliminating the reminiscence entry overhead of the 16-bit department. Determine from the SVDQuant paper.<\/figcaption><\/figure>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tIntroducing Nunchaku Lite<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>The unique Nunchaku engine will get a lot of its pace from model-specific fused execution paths, reminiscent of fused QKV projections and fused GELU\/MLP kernels. These optimizations are tied to every structure&#8217;s module format and checkpoint format, so supporting a brand new mannequin household normally requires model-specific integration work.<\/p>\n<p>Nunchaku Lite is the brand new integration path in Diffusers. With it, Diffusers can load Nunchaku-style checkpoints with no customized pipeline or a separate inference engine. Below the hood, Nunchaku Lite patches the related nn.Linear modules of a inventory Diffusers mannequin with runtime SVDQ\/AWQ linear layers earlier than the checkpoint is loaded. The CUDA kernels come from the Hub via the kernels package deal. Two kernel households are used:<\/p>\n<p>svdq_w4a4: 4-bit weights and activations with the SVDQuant low-rank correction. This layer is used for the transformer&#8217;s consideration and MLP projections, the place practically the entire compute is spent, and is accessible in INT4 and NVFP4 variants.<br \/>\nawq_w4a16: 4-bit weights with 16-bit activations, used for adaptive normalization and modulation projections reminiscent of FLUX adanorm_single \/ adanorm_zero or Qwen-Picture modulation layers. These layers are memory-bound and precision-sensitive, making AWQ an excellent match to protect precision whereas nonetheless saving reminiscence and house.<\/p>\n<p>The trade-off is that, with out architecture-specific fused kernels and modules, Nunchaku Lite can&#8217;t match the speedup of the unique Nunchaku engine. Nonetheless, the bare-bones implementation nonetheless delivers round 30% speedup whereas retaining the identical stage of VRAM discount.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tNative loading in Diffusers<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>When you have used bitsandbytes or torchao in Diffusers, the mechanics will really feel acquainted. A Nunchaku Lite mannequin repository is an peculiar Diffusers repository. The one particular half is a quantization_config block contained in the transformer&#8217;s config.json:<\/p>\n<p><span class=\"hljs-attr\">&#8220;quantization_config&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-punctuation\">{<\/span><br \/>\n    <span class=\"hljs-attr\">&#8220;quant_method&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-string\">&#8220;nunchaku_lite&#8221;<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n    <span class=\"hljs-attr\">&#8220;compute_dtype&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-string\">&#8220;bfloat16&#8221;<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n    <span class=\"hljs-attr\">&#8220;svdq_w4a4&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-punctuation\">{<\/span><br \/>\n        <span class=\"hljs-attr\">&#8220;precision&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-string\">&#8220;nvfp4&#8221;<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n        <span class=\"hljs-attr\">&#8220;group_size&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-number\">16<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n        <span class=\"hljs-attr\">&#8220;rank&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-number\">32<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n        <span class=\"hljs-attr\">&#8220;targets&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-punctuation\">[<\/span><br \/>\n            <span class=\"hljs-string\">&#8220;layers.0.self_attention.to_q&#8221;<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n            <span class=\"hljs-string\">&#8220;layers.0.self_attention.to_k&#8221;<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n            <span class=\"hljs-string\">&#8220;&#8230;&#8221;<\/span><br \/>\n        <span class=\"hljs-punctuation\">]<\/span><br \/>\n    <span class=\"hljs-punctuation\">}<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n    <span class=\"hljs-attr\">&#8220;awq_w4a16&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-punctuation\">{<\/span><br \/>\n        <span class=\"hljs-attr\">&#8220;precision&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-string\">&#8220;int4&#8221;<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n        <span class=\"hljs-attr\">&#8220;group_size&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-number\">64<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n        <span class=\"hljs-attr\">&#8220;targets&#8221;<\/span><span class=\"hljs-punctuation\">:<\/span> <span class=\"hljs-punctuation\">[<\/span><br \/>\n            <span class=\"hljs-string\">&#8220;adaLN_modulation.1&#8221;<\/span><span class=\"hljs-punctuation\">,<\/span><br \/>\n            <span class=\"hljs-string\">&#8220;&#8230;&#8221;<\/span><br \/>\n        <span class=\"hljs-punctuation\">]<\/span><br \/>\n    <span class=\"hljs-punctuation\">}<\/span><br \/>\n<span class=\"hljs-punctuation\">}<\/span><\/p>\n<p>This config tells Diffusers which modules have been quantized, which scheme they use, and which Nunchaku Lite runtime layer to instantiate (SVDQW4A4Linear or AWQW4A16Linear).<\/p>\n<p>As a result of the quantized mannequin retains the precise module construction of the dense one, the whole lot downstream (schedulers, LoRA loading hooks, offloading, torch.compile) sees a traditional Diffusers mannequin.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\t{Hardware} help<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>Nunchaku Lite makes use of completely different kernel variants relying on the GPU era and checkpoint precision:<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Scheme<br \/>\nPrecision<br \/>\nSupported GPUs<\/p>\n<p>svdq_w4a4<br \/>\nnvfp4<br \/>\nBlackwell (RTX 50 sequence, RTX PRO 6000, B200)<\/p>\n<p>svdq_w4a4<br \/>\nint4<br \/>\nTuring \/ Ampere \/ Ada (RTX 30 &amp; 40 sequence, A100, L40S)<\/p>\n<p>awq_w4a16<br \/>\nint4<br \/>\nTuring \/ Ampere \/ Ada (RTX 30 &amp; 40 sequence, A100, L40S)<\/p>\n<\/div>\n<blockquote class=\"warning\">\n<p>Volta and Hopper GPUs are presently not supported by the 4-bit kernels. The quantizer validates the GPU&#8217;s CUDA functionality at load time and raises a transparent error as an alternative of manufacturing incorrect outputs.<\/p>\n<\/blockquote>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tGetting extra pace and decrease reminiscence<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Nunchaku Lite could be mixed with different Diffusers reminiscence and pace optimizations.<\/p>\n<p>torch.compile. Compiling the transformer improves the end-to-end speedup from 1.35x to 1.8x:<\/p>\n<p>pipe.transformer.<span class=\"hljs-built_in\">compile<\/span>(fullgraph=<span class=\"hljs-literal\">True<\/span>)<\/p>\n<p>pipe.transformer.compile_repeated_blocks(fullgraph=<span class=\"hljs-literal\">True<\/span>)<\/p>\n<p>Quantized textual content encoders. The transformer will not be the one part with a big reminiscence footprint. Textual content encoders reminiscent of T5 or Qwen3 can occupy a number of gigabytes on their very own. Additional quantizing the textual content encoder with bitsandbytes NF4 reduces peak VRAM by about 22% in our benchmark.<\/p>\n<p>Offloading. Diffusers offloading helpers reminiscent of enable_model_cpu_offload() and enable_sequential_cpu_offload() work as ordinary if it&#8217;s essential match the pipeline onto a smaller GPU.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tBenchmarks<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>All numbers under have been measured on an NVIDIA RTX PRO 6000 (Blackwell) at 1024&#215;1024 utilizing rootonchair\/ERNIE-Picture-Turbo-nunchaku-lite-int4-bnb4-text-encoder.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tFinish-to-end latency and reminiscence<br \/>\n\t<\/span><br \/>\n<\/h3>\n<div class=\"max-w-full overflow-auto\">\n<p>Configuration<br \/>\nFull pipeline<br \/>\nDenoise loop<br \/>\nPeak VRAM<br \/>\nSpeedup<\/p>\n<p>BF16 baseline<br \/>\n3.00 s<br \/>\n2.86 s<br \/>\n31.1 GB<br \/>\n1.0x<\/p>\n<p>Nunchaku Lite NVFP4<br \/>\n2.27 s<br \/>\n2.13 s<br \/>\n20.6 GB<br \/>\n1.35x<\/p>\n<p>Nunchaku Lite NVFP4 + torch.compile<br \/>\n1.68 s<br \/>\n1.53 s<br \/>\n20.6 GB<br \/>\n1.8x<\/p>\n<p>Nunchaku Lite NVFP4 + NF4 textual content encoder<br \/>\n2.29 s<br \/>\n2.13 s<br \/>\n16.0 GB<br \/>\n1.35x<\/p>\n<\/div>\n<p>As proven above, Nunchaku reduces peak VRAM by as much as 50% whereas nonetheless bettering latency by roughly 30%. The remaining overhead comes largely from additional kernel launches, which torch.compile can mitigate, bringing the complete pipeline right down to 1.68 s, or 1.8x sooner than the BF16 baseline.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tPicture high quality<br \/>\n\t<\/span><br \/>\n<\/h3>\n<figure class=\"image text-center\">\n  <img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/nunchaku-diffusers\/quality_grid.png\" alt=\"Quality comparison grid\"\/><figcaption>BF16 vs 4-bit outputs with equivalent seeds and settings.<\/figcaption><\/figure>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tQuantizing your individual mannequin<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Nunchaku Lite help in Diffusers is architecture-agnostic, and the diffuse-compressor toolkit offers an end-to-end SVDQuant workflow for Diffusers fashions: calibrate, quantize, package deal, and publish.<\/p>\n<p>Beneath, we stroll via quantizing FLUX.2 Klein 4B for instance. It covers the principle steps: examine the mannequin, calibrate and quantize the transformer, package deal the end result as a Diffusers pipeline, then confirm and push it to the Hub. The complete tutorial covers each flag intimately.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\t1. Examine what will probably be quantized<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>The generic scanner walks the mannequin and decides what to focus on: appropriate linears contained in the repeated transformer-block stack turn out to be SVDQ W4A4 targets, acknowledged modulation linears turn out to be AWQ W4A16 targets, and the whole lot else stays dense.<\/p>\n<p>python examples\/text_to_image\/quantize_hf.py black-forest-labs\/FLUX.2-klein-4B<br \/>\n  &#8211;precision int4 &#8211;rank 32 &#8211;inspect-config<\/p>\n<p>All the time learn this report earlier than quantizing. For FLUX.2 Klein 4B, the anticipated result&#8217;s 100 SVDQ targets, 3 AWQ targets, and 6 dense outer linears, with no lacking patterns or duplicate names.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\t2. Run quantization<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>The next command runs SVDQuant on the transformer and writes the quantized checkpoint to outputs\/checkpoints\/svdq-int4_r32-flux-2-klein-4b.safetensors:<\/p>\n<p>python examples\/text_to_image\/quantize_hf.py black-forest-labs\/FLUX.2-klein-4B<br \/>\n  &#8211;precision int4<br \/>\n  &#8211;output outputs\/checkpoints\/svdq-int4_r32-flux-2-klein-4b.safetensors<\/p>\n<p>Substitute &#8211;precision int4 with nvfp4 to construct Blackwell-native weights.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\t3. Package deal a Diffusers pipeline<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>The converter combines the quantized transformer with the bottom pipeline&#8217;s different elements, writes the compact nunchaku_lite configuration into transformer\/config.json, and might optionally convert textual content encoders to NF4:<\/p>\n<p>python examples\/convert_nunchaku_lite_diffusers.py<br \/>\n  &#8211;checkpoint outputs\/checkpoints\/svdq-int4_r32-flux-2-klein-4b.safetensors<br \/>\n  &#8211;model-id black-forest-labs\/FLUX.2-klein-4B<br \/>\n  &#8211;bnb4-text-encoder text_encoder<br \/>\n  &#8211;compute-dtype bfloat16<br \/>\n  &#8211;output-dir outputs\/diffusers\/FLUX.2-klein-4B-nunchaku-lite-int4-bnb4-text-encoder<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\t4. Load, confirm, and push to the Hub<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p><span class=\"hljs-keyword\">import<\/span> torch<br \/>\n<span class=\"hljs-keyword\">from<\/span> diffusers <span class=\"hljs-keyword\">import<\/span> DiffusionPipeline<\/p>\n<p>pipe = DiffusionPipeline.from_pretrained(<br \/>\n    <span class=\"hljs-string\">&#8220;outputs\/diffusers\/FLUX.2-klein-4B-nunchaku-lite-int4-bnb4-text-encoder&#8221;<\/span>,<br \/>\n    device_map=<span class=\"hljs-string\">&#8220;cuda&#8221;<\/span>,<br \/>\n)<br \/>\npicture = pipe(<br \/>\n    <span class=\"hljs-string\">&#8220;A glass robotic in a greenhouse, cinematic lighting&#8221;<\/span>,<br \/>\n    num_inference_steps=<span class=\"hljs-number\">4<\/span>, guidance_scale=<span class=\"hljs-number\">1.0<\/span>,<br \/>\n    generator=torch.Generator(<span class=\"hljs-string\">&#8220;cuda&#8221;<\/span>).manual_seed(<span class=\"hljs-number\">12345<\/span>),<br \/>\n).photographs[<span class=\"hljs-number\">0<\/span>]<\/p>\n<p>As soon as the outputs look good, run pipe.push_to_hub(&#8220;your-name\/your-model-nunchaku-lite-int4&#8221;). Different customers can then load it with the identical from_pretrained() sample proven above.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tQuantizing fashions with structural rewrites<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>Be aware that the generic path assumes the structure could be quantized with out structural rewrites. For added speedup, the unique Nunchaku engine rewrites teams of Diffusers layers as fused modules. The generic path can&#8217;t infer these adjustments by itself, reminiscent of combining separate Q, Ok, and V projections into one module or splitting a fused projection throughout a number of modules.<\/p>\n<p>FLUX.1-dev&#8217;s QKV projection is a concrete instance. Diffusers defines three separate modules:<\/p>\n<p>self.to_q = torch.nn.Linear(query_dim, self.inner_dim, bias=bias)<br \/>\nself.to_k = torch.nn.Linear(query_dim, self.inner_dim, bias=bias)<br \/>\nself.to_v = torch.nn.Linear(query_dim, self.inner_dim, bias=bias)<\/p>\n<p>The Nunchaku FLUX module combines these layers into one quantized to_qkv module:<\/p>\n<p>to_qkv = fuse_linears([other.to_q, other.to_k, other.to_v])<br \/>\nself.to_qkv = SVDQW4A4Linear.from_linear(to_qkv, **kwargs)<\/p>\n<p>This grouped module is required as a result of Nunchaku&#8217;s fused operator consumes the QKV projection, Q\/Ok normalization, and rotary embeddings collectively. By comparability, the default Diffusers path executes them individually:<\/p>\n<p>question = attn.to_q(hidden_states)<br \/>\nkey = attn.to_k(hidden_states)<br \/>\nworth = attn.to_v(hidden_states)<\/p>\n<p>question = question.unflatten(-<span class=\"hljs-number\">1<\/span>, (attn.heads, &#8211;<span class=\"hljs-number\">1<\/span>))<br \/>\nkey = key.unflatten(-<span class=\"hljs-number\">1<\/span>, (attn.heads, &#8211;<span class=\"hljs-number\">1<\/span>))<br \/>\nworth = worth.unflatten(-<span class=\"hljs-number\">1<\/span>, (attn.heads, &#8211;<span class=\"hljs-number\">1<\/span>))<\/p>\n<p>question = attn.norm_q(question)<br \/>\nkey = attn.norm_k(key)<\/p>\n<p><span class=\"hljs-keyword\">if<\/span> image_rotary_emb <span class=\"hljs-keyword\">is<\/span> <span class=\"hljs-keyword\">not<\/span> <span class=\"hljs-literal\">None<\/span>:<br \/>\n    question = apply_rotary_emb(question, image_rotary_emb, sequence_dim=<span class=\"hljs-number\">1<\/span>)<br \/>\n    key = apply_rotary_emb(key, image_rotary_emb, sequence_dim=<span class=\"hljs-number\">1<\/span>)<\/p>\n<p>The Nunchaku path provides the grouped projection, normalization modules, and rotary embeddings to at least one fused operator:<\/p>\n<p>qkv = fused_qkv_norm_rottary(<br \/>\n    hidden_states, attn.to_qkv, attn.norm_q, attn.norm_k, image_rotary_emb<br \/>\n)<\/p>\n<p>That is the structural rewrite that the generic path can&#8217;t infer. Diffusers has three vacation spot modules with to_q, to_k, and to_v parameter prefixes, whereas Nunchaku has one grouped module beneath to_qkv. A model-specific goal config or adapter should state that the Q, Ok, and V parameters needs to be concatenated alongside the output dimension, in that order, and loaded into to_qkv.<\/p>\n<p>Structural rewrites like these are described by a model-specific goal config throughout quantization and dealt with by a small runtime adapter when the checkpoint is loaded.<br \/>\nThe FLUX.2 Klein 4B quantization script offers a concrete target-config instance for producing a structurally rewritten checkpoint, whereas rootonchair\/nunchaku-lite offers the runtime adapters wanted to load grouped QKV tensors, cut up fused projections, and different fused operations.<br \/>\nFor the whole workflow, you&#8217;ll be able to verify the Including A New Mannequin information.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tPrepared-to-use checkpoints<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>To get began instantly, take a look at the next repositories:<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tConclusion<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Nunchaku&#8217;s SVDQuant kernels are some of the efficient methods to run diffusion transformers effectively on client {hardware}, and they&#8217;re now natively supported in Diffusers. Pre-quantized checkpoints load with from_pretrained(), and the diffuse-compressor toolkit makes it attainable to quantize new architectures with out ready for engine help. By quantizing each weights and activations, the W4A4 path lowers reminiscence use whereas bettering denoising latency, protecting picture high quality near the BF16 authentic.<\/p>\n<p>If you happen to quantize and publish a brand new mannequin, we might love to listen to about it. Share it on the Hub and tell us! When you have any questions on this characteristic, be happy to affix our Discord.<\/p>\n<p>To study extra, take a look at the next assets:<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tAcknowledgements<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Due to the Diffusers maintainers for evaluations and steering all through the combination, and to the MIT HAN Lab \/ Nunchaku staff for the unique SVDQuant work. Due to Marc Solar for offering suggestions on the weblog publish. Due to \u00c1lvaro Somoza for attempting out nunchaku-lite and for offering suggestions.<\/p>\n<p>rootonchair can be grateful to SilverAI for supporting this work and offering the setting through which a lot of this improvement passed off.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/huggingface.co\/blog\/nunchaku-diffusers\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Massive diffusion transformers can create gorgeous photographs (and even movies, audio snippets, and now textual content), however loading a contemporary text-to-image mannequin in BF16 precision typically requires 20-30 GB of VRAM, which places these fashions out of attain of most client GPUs. Quantization is a robust resolution to this downside, and Diffusers already integrates a [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":2788,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/huggingface.co\/blog\/assets\/nunchaku-diffusers\/thumbnail.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[5],"tags":[3252,3250,3076,2389,1068,3251],"class_list":["post-2786","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-developer-ai-open-source-ecosystem","tag-4bit","tag-bringing","tag-diffusers","tag-diffusion","tag-inference","tag-nunchaku"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Bringing Nunchaku 4-bit Diffusion Inference to Diffusers - Future News 24<\/title>\n<meta name=\"description\" content=\"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers - Future News 24\" \/>\n<meta property=\"og:description\" content=\"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-23T00:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-24T12:59:07+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/huggingface.co\/blog\/assets\/nunchaku-diffusers\/thumbnail.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/huggingface.co\/blog\/assets\/nunchaku-diffusers\/thumbnail.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"12 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers\",\"datePublished\":\"2026-07-23T00:00:00+00:00\",\"dateModified\":\"2026-07-24T12:59:07+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/\"},\"wordCount\":2359,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/huggingface.co\\\/blog\\\/assets\\\/nunchaku-diffusers\\\/thumbnail.png\",\"keywords\":[\"4bit\",\"Bringing\",\"Diffusers\",\"Diffusion\",\"inference\",\"Nunchaku\"],\"articleSection\":[\"Developer AI &amp; Open-Source Ecosystem\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/\",\"name\":\"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/huggingface.co\\\/blog\\\/assets\\\/nunchaku-diffusers\\\/thumbnail.png\",\"datePublished\":\"2026-07-23T00:00:00+00:00\",\"dateModified\":\"2026-07-24T12:59:07+00:00\",\"description\":\"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/#primaryimage\",\"url\":\"https:\\\/\\\/huggingface.co\\\/blog\\\/assets\\\/nunchaku-diffusers\\\/thumbnail.png\",\"contentUrl\":\"https:\\\/\\\/huggingface.co\\\/blog\\\/assets\\\/nunchaku-diffusers\\\/thumbnail.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/23\\\/nunchaku-diffusers\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers - Future News 24","description":"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/","og_locale":"en_US","og_type":"article","og_title":"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers - Future News 24","og_description":"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/","og_site_name":"Future News 24","article_published_time":"2026-07-23T00:00:00+00:00","article_modified_time":"2026-07-24T12:59:07+00:00","og_image":[{"url":"https:\/\/huggingface.co\/blog\/assets\/nunchaku-diffusers\/thumbnail.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/huggingface.co\/blog\/assets\/nunchaku-diffusers\/thumbnail.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"12 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers","datePublished":"2026-07-23T00:00:00+00:00","dateModified":"2026-07-24T12:59:07+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/"},"wordCount":2359,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/#primaryimage"},"thumbnailUrl":"https:\/\/huggingface.co\/blog\/assets\/nunchaku-diffusers\/thumbnail.png","keywords":["4bit","Bringing","Diffusers","Diffusion","inference","Nunchaku"],"articleSection":["Developer AI &amp; Open-Source Ecosystem"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/","name":"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/#primaryimage"},"thumbnailUrl":"https:\/\/huggingface.co\/blog\/assets\/nunchaku-diffusers\/thumbnail.png","datePublished":"2026-07-23T00:00:00+00:00","dateModified":"2026-07-24T12:59:07+00:00","description":"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/#primaryimage","url":"https:\/\/huggingface.co\/blog\/assets\/nunchaku-diffusers\/thumbnail.png","contentUrl":"https:\/\/huggingface.co\/blog\/assets\/nunchaku-diffusers\/thumbnail.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/23\/nunchaku-diffusers\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Bringing Nunchaku 4-bit Diffusion Inference to Diffusers"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2786","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=2786"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2786\/revisions"}],"predecessor-version":[{"id":2787,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2786\/revisions\/2787"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/2788"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=2786"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=2786"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=2786"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}