{"id":3146,"date":"2026-07-31T22:16:00","date_gmt":"2026-07-31T22:16:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/"},"modified":"2026-08-01T17:59:46","modified_gmt":"2026-08-01T17:59:46","slug":"co-designing-ai-model-attention-for-fast-interactive-long-context-inference","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/","title":{"rendered":"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">As agentic and long-context workloads grow to be widespread, the context lengths improve and a spotlight consumes a bigger share of inference time (Determine 1). As a result of consideration now dominates that price, how it&#8217;s designed\u2014not simply how it&#8217;s applied\u2014more and more determines a mannequin\u2019s inference efficiency. Shaping mannequin structure round how GPUs execute it&#8217;s the premise of AI mannequin co-design. For a dialogue of how mannequin design selections affect each throughput and interactivity with out sacrificing accuracy, see the earlier submit, AI Mannequin Co-Design: {Hardware}-Pleasant LLM Design.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">This submit examines how group measurement (question heads per KV head), head dimension, and sequence size form the efficiency of dense consideration, the place each question attends to all keys and values alongside the sequence size. We distill that evaluation, along with how consideration is parallelized throughout GPUs, into 4 sensible pointers: a co-design guidelines that helps mannequin builders elevate inference throughput and interactivity on NVIDIA GPUs. Keep tuned for a submit protecting sparse consideration.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a6e34114dc6d&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a6e34114dc6d\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"936\" height=\"298\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model.webp\" alt=\"Pie charts showing the attention prefill share increasing from 18% to 85% as context grows from 4K to 128K tokens. &#10;\" class=\"wp-image-120735\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model.webp 936w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-179x57.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-300x96.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-768x245.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-625x199.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-645x205.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-500x159.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-160x51.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-362x115.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-346x110.png 346w\" sizes=\"(max-width: 936px) 100vw, 936px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"936\" height=\"298\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model.webp\" alt=\"Pie charts showing the attention prefill share increasing from 18% to 85% as context grows from 4K to 128K tokens. &#10;\" class=\"lazyload wp-image-120735\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model.webp 936w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-179x57.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-300x96.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-768x245.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-625x199.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-645x205.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-500x159.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-160x51.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-362x115.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/attention-prefill-time-ai-model-346x110.png 346w\" data-sizes=\"(max-width: 936px) 100vw, 936px\"\/><figcaption class=\"wp-element-caption\">Determine 1. DeepSeek-R1 prefill time breakdown at 4K, 32K, and 128K context lengths, the place the eye share rises from 18% to 85%<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Each evaluation is grounded in two sources: analytical formulation from GEMM-shape arithmetic, and measured knowledge from prefill and decode kernels with FP8 for each consideration compute and the KV cache.\u00a0<\/p>\n<figure class=\"wp-block-table aligncenter\">AbbreviationDefinition\u00a0PB\u00a0Prefill batch measurement\u00a0DB\u00a0Decode batch measurement\u00a0QH\u00a0Variety of question heads\u00a0KH\u00a0Variety of KV heads (KH = QH for MHA, KH = QH\/G for GQA, KH = 1 for MQA)\u00a0(G)\u00a0Group measurement = QH\/KH (question heads sharing one KV head)\u00a0Hsz\u00a0Head dimension (sometimes 64, 128, or 256)\u00a0ISL\u00a0Enter sequence size (question tokens in prefill)\u00a0KVSL\u00a0Common KV cache sequence size in a decode iteration\u00a0<figcaption class=\"wp-element-caption\">Desk 1. Notations used within the equations featured on this submit\u00a0<\/figcaption><\/figure>\n<h2 id=\"how_are_prefill_and_decode_two_different_problems\u00a0\" class=\"wp-block-heading\">How are prefill and decode two completely different issues?\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Prefill processes the complete immediate in parallel, producing massive GEMM-M (= ISL \u00d7 (G)) matmuls which are compute-bound. With out speculative decoding, decode generates one token at a time, producing small GEMM-M (= (G)) matmuls and turning into memory-bound by KV cache reads from high-bandwidth reminiscence (HBM).\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Speculative decoding will increase GEMM-M and may shift decode towards compute-bound. As a result of the question lengths, KV entry, and bottleneck differ for prefill and decode (Desk 2), every parameter is analyzed individually for every part. \u00a0<\/p>\n<figure class=\"wp-block-table aligncenter\">\u00a0Prefill\u00a0Decode\u00a0Question size\u00a0Full enter (ISL tokens)\u00a01 token\u00a0KV context\u00a0Immediate (ISL tokens)\u00a0Full KV cache (KVSL tokens)\u00a0Consideration GEMM-M\u00a0ISL \u00d7 (G) (massive)\u00a0(G) (small)\u00a0Major bottleneck\u00a0Compute (matmul + softmax)\u00a0HBM bandwidth (reminiscence)\u00a0<figcaption class=\"wp-element-caption\">Desk 2. Prefill and decode differ in question size, KV context, consideration GEMM-M, and first bottleneck<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">Be aware: With prefix caching, widespread in agentic and multiturn apps, a brand new flip could have a brief ISL whereas attending to a big prefix cache. With a brief ISL however lengthy prefix cache, prefill behaves like decode.<\/p>\n<h2 id=\"how_arithmetic_intensity_governs_compute-_versus_memory-bound_behavior\" class=\"wp-block-heading\">How arithmetic depth governs compute- versus memory-bound habits<\/h2>\n<p class=\"wp-block-paragraph\">The roofline bounds GPU efficiency by compute and bandwidth ceilings, as beforehand defined. Arithmetic depth determines which binds (Equation 1):\u00a0<\/p>\n<p class=\"has-text-align-center wp-block-paragraph\">Arithmetic Depth = Whole FLOPs \/ Whole bytes accessed<\/p>\n<p class=\"wp-block-paragraph\">The ridge level marks the transition from memory-bound to compute-bound. Prefill lies properly above it and is compute-bound, whereas decode lies under it and is memory-bound (Determine 2). Speculative decoding raises decode arithmetic depth and may transfer it towards the ridge.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a6e34114edf4&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a6e34114edf4\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"454\" height=\"328\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode.webp\" alt=\"Roofline plot showing prefill above the ridge point and decode below it. \" class=\"wp-image-120743\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode.webp 454w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-159x115.png 159w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-300x217.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-415x300.png 415w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-125x90.png 125w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-362x262.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-152x110.png 152w\" sizes=\"(max-width: 454px) 100vw, 454px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"454\" height=\"328\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode.webp\" alt=\"Roofline plot showing prefill above the ridge point and decode below it. \" class=\"lazyload wp-image-120743\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode.webp 454w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-159x115.png 159w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-300x217.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-415x300.png 415w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-125x90.png 125w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-362x262.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/roofline-plot-prefill-decode-152x110.png 152w\" data-sizes=\"(max-width: 454px) 100vw, 454px\"\/><figcaption class=\"wp-element-caption\">Determine 2. Roofline mannequin exhibiting decode on the memory-bound ramp and prefill on the compute-bound plateau\u00a0<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"how_does_the_flashattention_kernel_compute_attention_on_gpu\u00a0\" class=\"wp-block-heading\">How does the FlashAttention kernel compute consideration on GPU?\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">FlashAttention computes consideration with out materializing the complete consideration matrix. It streams tiles of (Q), (Ok), and (V) from HBM to on-chip SRAM and fuses three steps into one cross:\u00a0<\/p>\n<p>First, batched matmul (BMM1) scores queries in opposition to keys<\/p>\n<p>Second, on-line softmax normalizes the scores utilizing a working max and sum<\/p>\n<p>Third, second batched matmul (BMM2) weights the values<\/p>\n<p class=\"wp-block-paragraph\">The BMMs run on Tensor Cores whereas the softmax exponentials run on special-function items. The BMM shapes drive the arithmetic depth evaluation that follows.\u00a0<\/p>\n<h2 id=\"gemm_shapes\u00a0\" class=\"wp-block-heading\">GEMM shapes\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Consideration efficiency follows from the shapes of its two matmuls. Desk 3 lists the per-phase (Batch, M, N, Ok) dimensions of BMM1 and BMM2. \u00a0<\/p>\n<figure class=\"wp-block-table aligncenter\">BMM\u00a0Section\u00a0Batch\u00a0M\u00a0N\u00a0Ok\u00a0Which means\u00a0BMM1\u00a0Prefill\u00a0PB \u00d7 KH\u00a0ISL \u00d7 (G)\u00a0ISL\u00a0Hsz\u00a0Q \u00b7 K\u1d40: rating queries in opposition to keys\u00a0Decode\u00a0DB \u00d7 KH\u00a01 \u00d7 (G)\u00a0KVSL\u00a0Hsz\u00a0BMM2\u00a0Prefill\u00a0PB \u00d7 KH\u00a0ISL \u00d7 (G)\u00a0Hsz\u00a0ISL\u00a0Weights \u00b7 V: combination values\u00a0Decode\u00a0DB \u00d7 KH\u00a01 \u00d7 (G)\u00a0Hsz\u00a0KVSL\u00a0<figcaption class=\"wp-element-caption\">Desk 3. GEMM (Batch, M, N, Ok) shapes for BMM1 and BMM2<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">For decode, GEMM-M = (G), normally 8-16, far under a GPU tile-M of 64 or 128, limiting parallel work per tile. Bigger (G) masses much less KV per token and amortizes every load throughout extra question heads, enhancing utilization. The following part quantifies this impact.\u00a0<\/p>\n<h2 id=\"group_size\u00a0\" class=\"wp-block-heading\">Group measurement\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Group measurement ((G)) is the variety of question heads that share one KV head. MHA has (G) = 1, GQA has (G) = 4, 8, 16, \u2026, and MQA has (G) = QH.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Arithmetic depth as a perform of (G). Within the following formulation, \u201cBytes\u201d refers to HBM bytes moved. For simplicity, assume 1 byte per ingredient (that&#8217;s, FP8 KV cache). \u00a0<\/p>\n<h3 id=\"prefill\u00a0\" class=\"wp-block-heading\">Prefill\u00a0<\/h3>\n<p class=\"wp-block-paragraph\">As (G) grows, the 1\/(G) vanishes and arithmetic depth approaches 2 \u00d7 ISL. At ISL = 32K, elevating (G) from 8 to 16 improves arithmetic depth by beneath 6%. In different phrases, prefill is dominated by ISL, not (G). Determine 4 confirms this, various (G) from 1 (MHA) to 64 (MQA) adjustments prefill runtime by beneath 1%. Equations 2, 3, and 4:<\/p>\n<p class=\"has-text-align-center wp-block-paragraph\">FLOPs = 4 \u00d7 PB \u00d7 QH \u00d7 ISL\u00b2 \u00d7 Hsz (fixed in (G))Bytes = 2 \u00d7 PB \u00d7 KH \u00d7 Hsz \u00d7 ISL \u00d7 ((G) + 1)Arithmetic Depth = 2 \u00d7 (G) \u00d7 ISL \/ ((G) + 1) = 2 \u00d7 ISL \/ (1 + 1\/(G)) \u2192 2 \u00d7 ISL as (G) \u2192 \u221e\u00a0 \u00a0 \u00a0<\/p>\n<h3 id=\"decode_gemm-m_=_g\" class=\"wp-block-heading\">Decode (GEMM-M = (G))<\/h3>\n<p class=\"wp-block-paragraph\">Doubling (G) doubles decode arithmetic depth. Elevating (G) from 1 to eight offers an 8x achieve by decreasing reminiscence visitors and enhancing GPU compute utilization. It&#8217;s unbiased of KVSL: arithmetic depth stays close to 2 \u00d7 (G), so decode stays memory-bound until (G) may be very massive. Fashions akin to NVIDIA Nemotron 3 adopted GQA with two KV heads, which makes decode extra environment friendly. Equations 5, 6, and seven:<\/p>\n<p class=\"has-text-align-center wp-block-paragraph\">FLOPs = 4 \u00d7 DB \u00d7 QH \u00d7 KVSL \u00d7 Hsz (fixed in (G)) \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0Bytes = 2 \u00d7 DB \u00d7 KH \u00d7 Hsz \u00d7 ((G) + KVSL)Arithmetic Depth = 2 \u00d7 (G) \u00d7 KVSL \/ ((G) + KVSL) \u2248 2 \u00d7 (G) (when KVSL \u226b (G))<\/p>\n<p class=\"wp-block-paragraph\">Determine 4 exhibits decode runtime falls about 2x per doubling of (G) as a result of halving the KV heads halves the info loaded per token. Past (G) = 16, the KVSL = 32K curve flattens. Its per-step kernel is sufficiently small that two prices dominate: mounted setup and post-processing overhead, and the flash-decoding discount from splitting KV throughout SMs to remain parallel with few KV heads. The longer KVSL = 128K kernel higher amortizes these prices and continues monitoring the 2x pattern.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Be aware: Speculative decoding raises efficient GEMM-M to (1+(D)) \u00d7 (G), the place (D) is the variety of draft tokens. As soon as massive sufficient to fill compute tiles, decode strikes towards compute-bound.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a6e341150ef6&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a6e341150ef6\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"933\" height=\"381\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g.webp\" alt=\"Two line graphs showing runtime versus G. Prefill is flat; decode falls about 2\u00d7 per doubling of G.&#10;\" class=\"wp-image-120749\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g.webp 933w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-179x73.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-300x123.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-768x314.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-625x255.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-645x263.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-500x204.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-160x65.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-362x148.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-269x110.png 269w\" sizes=\"(max-width: 933px) 100vw, 933px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"933\" height=\"381\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g.webp\" alt=\"Two line graphs showing runtime versus G. Prefill is flat; decode falls about 2\u00d7 per doubling of G.&#10;\" class=\"lazyload wp-image-120749\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g.webp 933w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-179x73.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-300x123.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-768x314.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-625x255.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-645x263.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-500x204.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-160x65.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-362x148.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-g-269x110.png 269w\" data-sizes=\"(max-width: 933px) 100vw, 933px\"\/><figcaption class=\"wp-element-caption\">Determine 4. Normalized runtime versus (G) (QH=64, Hsz=128, PB=1, DB=8; prefill ISL and decode KVSL as 32K and 128K). Prefill is flat in (G) (compute-bound); decode falls ~2x per doubling of (G) (memory-bound), from MHA ((G)=1) to MQA ((G)=64)<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Guideline 1: Select (G) for decode effectivity and push it excessive. Prefill runtime is flat in (G), whereas decode arithmetic depth \u2248 2 \u00d7 (G), so greater (G) improves decode velocity and GPU utilization. Speculative decoding is one other lever to spice up efficiency at a given (G).<\/p>\n<h2 id=\"head_dimension\u00a0\u00a0\" class=\"wp-block-heading\">Head dimension\u00a0\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Not like group measurement, head dimension (Hsz) doesn&#8217;t have an effect on arithmetic depth. Doubling Hsz doubles each FLOPs (Equations 2 and 5) and bytes (Equations 3 and 6), leaving their ratio unchanged. But Determine 5 exhibits runtime growing with Hsz as a result of consideration kernels carry out three forms of work that scale in a different way with Hsz.\u00a0<\/p>\n<p>Matmul: Grows with Hsz, however in aligned steps. The earlier submit recommends mannequin dimensions which are multiples of 128 to align with GPU tile sizes and cache-line widths. {A partially} crammed tile prices as a lot as a full tile, so Hsz = 64 pays for 128. Hsz \u2265 512 pushes near the tensor reminiscence (TMEM) capability restrict. That makes 128 and 256 the environment friendly selections.\u00a0<\/p>\n<p>Reminiscence (the KV state): Additionally grows with Hsz. Because the GPU strikes knowledge in 128-byte items, reminiscence entry, like matmul, is most effective when Hsz is a number of of 128.\u00a0<\/p>\n<p>Softmax is completely different: Its price is unbiased of Hsz as a result of it operates on attention-score matrix (question tokens \u00d7 keys), which has no head dimension. Equations 8 and 9:\u00a0<\/p>\n<p class=\"has-text-align-center wp-block-paragraph\">Softmax ops (prefill) \u2248 PB \u00d7 QH \u00d7 ISL\u00b2Softmax ops (decode) \u2248 DB \u00d7 QH \u00d7 KVSL<\/p>\n<p class=\"wp-block-paragraph\">The stability between these determines Hsz price in every part (Determine 5).\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a6e341151dfa&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a6e341151dfa\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"933\" height=\"381\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz.webp\" alt=\"Two line graphs showing runtime versus Hsz. Runtime rises with Hsz for both prefill and decode.&#10;\" class=\"wp-image-120773\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz.webp 933w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-179x73.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-300x123.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-768x314.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-625x255.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-645x263.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-500x204.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-160x65.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-362x148.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-269x110.png 269w\" sizes=\"(max-width: 933px) 100vw, 933px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"933\" height=\"381\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz.webp\" alt=\"Two line graphs showing runtime versus Hsz. Runtime rises with Hsz for both prefill and decode.&#10;\" class=\"lazyload wp-image-120773\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz.webp 933w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-179x73.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-300x123.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-768x314.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-625x255.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-645x263.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-500x204.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-160x65.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-362x148.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-hsz-269x110.png 269w\" data-sizes=\"(max-width: 933px) 100vw, 933px\"\/><figcaption class=\"wp-element-caption\">Determine 5. Normalized runtime versus Hsz (QH=64, (G)=32, PB=1, DB=8; prefill ISL and decode KVSL as 32K and 128K). Runtime rises with Hsz\u00a0<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Prefill is compute-bound (matmul plus softmax). As Hsz grows, matmul FLOPs develop whereas softmax stays mounted. If prefill had been pure matmul, doubling Hsz would double runtime; however the mounted softmax doesn&#8217;t scale, so runtime rises lower than Hsz does. Determine 5 confirms this: prefill climbs with Hsz however slower than Hsz grows. A wider Hsz amortizes softmax, shifting extra kernel time to matmul and making prefill much less softmax-bound.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Decode is memory-bound (streaming the KV cache). Bigger Hsz will increase KV bytes per token (Equation 6), so runtime ought to scale with Hsz. Determine 5 confirms this, although barely sublinearly as a result of setup, post-processing, and flash-decoding discount overheads don\u2019t scale with Hsz and weigh extra on the shorter KVSL = 32K kernel.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Guideline 2: Use Hsz of 128 or 256. Hsz doesn\u2019t change arithmetic depth however should align with the {hardware}. Usually, Hsz = 64 nonetheless pays for a 128-wide tile, whereas Hsz \u2265 512 pushes near the TMEM capability restrict. That makes 128 and 256 the candy spot.\u00a0<\/p>\n<h2 id=\"sequence_length\u00a0\" class=\"wp-block-heading\">Sequence size\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Sequence size (ISL \/ KVSL) impacts prefill and decode in a different way, as a result of it enters every part by a special variable: prefill processes all ISL enter tokens collectively, whereas every decode step reads a KV cache of size KVSL. The 2 subsequently scale at completely different charges (Determine 6).\u00a0<\/p>\n<h3 id=\"prefill_scales_quadratically\u00a0\" class=\"wp-block-heading\">Prefill scales quadratically\u00a0<\/h3>\n<p class=\"wp-block-paragraph\">Prefill performs ISL\u00b2 work (each token attends to each token), whereas KV visitors grows solely proportional to ISL (Eq. 2, 3). Arithmetic depth subsequently rises linearly with ISL, maintaining prefill properly above the ridge level and compute-bound. Doubling ISL ought to roughly quadruple runtime (Determine 6). At quick ISL, scaling is under 4x as a result of mounted setup and post-processing overheads dominate; quadratic scaling seems as soon as ISL is massive sufficient to amortize them.\u00a0<\/p>\n<h3 id=\"decode_scales_linearly\u00a0\" class=\"wp-block-heading\">Decode scales linearly\u00a0<\/h3>\n<p class=\"wp-block-paragraph\">Every step reads the complete KV cache to generate one token, so bytes develop with KVSL whereas per-step work stays small (Equations 5 and 6). Arithmetic depth stays close to 2 \u00d7 (G), properly under the ridge level, so decode stays memory-bound in any respect lengths. Doubling KVSL ought to double runtime, which Determine 6 confirms. At quick KVSL, scaling is under 2x as a result of setup, post-processing, and flash-decoding discount overheads don\u2019t scale with KVSL. Their share shrinks as KVSL grows.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a6e341152f6d&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a6e341152f6d\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"936\" height=\"379\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length.webp\" alt=\" Two line graphs showing runtime versus sequence length; prefill scales as O(n\u00b2); decode as O(n). &#10;\" class=\"wp-image-120777\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length.webp 936w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-179x72.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-300x121.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-768x311.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-625x253.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-645x261.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-500x202.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-160x65.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-362x147.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-272x110.png 272w\" sizes=\"(max-width: 936px) 100vw, 936px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"936\" height=\"379\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length.webp\" alt=\" Two line graphs showing runtime versus sequence length; prefill scales as O(n\u00b2); decode as O(n). &#10;\" class=\"lazyload wp-image-120777\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length.webp 936w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-179x72.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-300x121.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-768x311.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-625x253.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-645x261.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-500x202.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-160x65.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-362x147.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/normalized-runtime-versus-sequence-length-272x110.png 272w\" data-sizes=\"(max-width: 936px) 100vw, 936px\"\/><figcaption class=\"wp-element-caption\">Determine 6. Normalized runtime versus sequence size (QH=64, (G)=32, Hsz=256, PB=1, DB=8). Prefill scales as O(n\u00b2) with ISL; decode as O(n) with KVSL; each scale under the perfect fee at quick lengths<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Guideline 3: Scale back efficient KV state the place doable. The use case units sequence size, however the associated fee is uneven: prefill grows as ISL\u00b2, whereas decode grows linearly with KVSL. Scale back efficient KV state by KV-cache compression, sparse or sliding-window consideration, or hybrid mannequin architectures (like Nemotron 3) the place just some layers carry the rising international KV state.\u00a0<\/p>\n<h2 id=\"tensor_parallelism_splits_attention_heads_across_gpus_\u00a0\" class=\"wp-block-heading\">Tensor parallelism splits consideration heads throughout GPUs \u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Tensor parallelism (TP) splits the eye heads throughout GPUs, giving every GPU QH\/TP question heads and KH\/TP KV heads. It shards heads, not tokens. In Desk 3, solely the batch dimension, containing KH\/TP, shrinks; the per-GPU GEMM form and arithmetic depth stay unchanged. \u00a0<\/p>\n<p class=\"wp-block-paragraph\">TP has a sensible restrict: KV heads should divide evenly throughout GPUs. As soon as TP &gt; KH, a bunch\u2019s question heads span a number of ranks, every requiring a replica of the shared KV head. This duplicates KV state, including reminiscence and bandwidth overhead with out profit (Determine 7). Thus, hold TP \u2264 KH so every GPU owns at the very least one full group: one KV head and its (G) question heads.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Fashions with few KV heads (Nemotron 3 with two, for instance), shortly exhaust TP as a result of the cache can&#8217;t be sharded under one KV head per GPU with out duplication. Consideration should then scale in a different way: Consideration Information Parallelism (ADP) shards requests, whereas KV Parallelism (KVP) shards long-sequence KV caches throughout GPUs. The FFN scales individually with Knowledgeable Parallelism (EP).\u00a0<\/p>\n<p class=\"wp-block-paragraph\">TensorRT-LLM combines these as Huge EP (ADP for consideration plus EP for the FFN), and Helix Parallelism (KVP for consideration plus EP for the FFN). In each circumstances, KH determines environment friendly scaling.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a6e341153dda&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a6e341153dda\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"847\" height=\"439\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies.webp\" alt=\"A diagram showing how tensor parallel (TP) works across GPUs for different settings\u2014(a) no TP, (b) TP=2, and (c) TP=4\u2014including how activations (V, K, Q) and resulting token outputs are distributed. It also highlights duplicated segments when the TP setting causes repeated computation or data movement.&#10;\" class=\"wp-image-120783\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies.webp 847w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-179x93.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-300x155.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-768x398.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-625x324.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-645x334.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-500x259.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-160x83.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-362x188.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-212x110.png 212w\" sizes=\"(max-width: 847px) 100vw, 847px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"847\" height=\"439\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies.webp\" alt=\"A diagram showing how tensor parallel (TP) works across GPUs for different settings\u2014(a) no TP, (b) TP=2, and (c) TP=4\u2014including how activations (V, K, Q) and resulting token outputs are distributed. It also highlights duplicated segments when the TP setting causes repeated computation or data movement.&#10;\" class=\"lazyload wp-image-120783\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies.webp 847w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-179x93.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-300x155.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-768x398.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-625x324.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-645x334.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-500x259.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-160x83.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-362x188.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/atention-sharding-strategies-212x110.png 212w\" data-sizes=\"(max-width: 847px) 100vw, 847px\"\/><figcaption class=\"wp-element-caption\">Determine 7. Consideration sharding methods: when TP &gt; KH, the KV cache is duplicated, including reminiscence and bandwidth overhead<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Guideline 4: Let KH decide the parallelism technique. Preserve TP \u2264 KH so every GPU has a full KV head. Fashions with few KV heads (MQA with 1, or GQA with 2) exhaust TP shortly and are higher served by ADP or KVP for consideration plus EP for the MoE FFN (applied as Huge EP and Helix Parallelism in TensorRT-LLM).\u00a0<\/p>\n<h2 id=\"get_started_co-designing_ai_model_attention\" class=\"wp-block-heading\">Get began co-designing AI mannequin consideration<\/h2>\n<p class=\"wp-block-paragraph\">Use the 4 pointers summarized under as a mannequin design guidelines to get began co-designing AI mannequin consideration. These selections can enhance GPU utilization, inference velocity, throughput, and interactivity on the identical {hardware}.\u00a0\u00a0<\/p>\n<p>Guideline 1: Select the group measurement ((G)) for decode and push it excessive. Prefill just isn&#8217;t delicate to group measurement.\u00a0<\/p>\n<p>Guideline 2: Use a head dimension (Hsz) = 128 or 256 to align with the GPU tiles and 128-byte transfers whereas staying inside TMEM finances. A bigger head additionally hides softmax in prefill.\u00a0<\/p>\n<p>Guideline 3: Scale back efficient KV state by KV-cache compression, sparse or sliding-window consideration, or hybrid fashions.\u00a0<\/p>\n<p>Guideline 4: Match parallelism to KV head depend (KH). Preserve TP \u2264 KH and scale a number of KH fashions with Huge EP and Helix Parallelism.\u00a0<\/p>\n<h3 id=\"acknowledgments\u00a0\" class=\"wp-block-heading\">Acknowledgments\u00a0<\/h3>\n<p class=\"wp-block-paragraph\">This submit is an NVIDIA cross-team effort. We&#8217;re grateful to Timmy Liu, Jatin Mitra, Tiyasa Mitra, Bhargava Gopireddy, Brian Pharris, Julien Demouth, and Eduardo Alvarez for his or her assist.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>As agentic and long-context workloads grow to be widespread, the context lengths improve and a spotlight consumes a bigger share of inference time (Determine 1). As a result of consideration now dominates that price, how it&#8217;s designed\u2014not simply how it&#8217;s applied\u2014more and more determines a mannequin\u2019s inference efficiency. Shaping mannequin structure round how GPUs execute [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3148,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/llm-optimize-deploy.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[1050,3563,417,1068,2040,757,105],"class_list":["post-3146","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-attention","tag-codesigning","tag-fast","tag-inference","tag-interactive","tag-longcontext","tag-model"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference - Future News 24<\/title>\n<meta name=\"description\" content=\"As agentic and long&#x2d;context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1).\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference - Future News 24\" \/>\n<meta property=\"og:description\" content=\"As agentic and long&#x2d;context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1).\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-31T22:16:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-01T17:59:46+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/llm-optimize-deploy.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/llm-optimize-deploy.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"12 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference\",\"datePublished\":\"2026-07-31T22:16:00+00:00\",\"dateModified\":\"2026-08-01T17:59:46+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/\"},\"wordCount\":2347,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/llm-optimize-deploy.webp\",\"keywords\":[\"Attention\",\"CoDesigning\",\"Fast\",\"inference\",\"interactive\",\"LongContext\",\"Model\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/\",\"name\":\"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/llm-optimize-deploy.webp\",\"datePublished\":\"2026-07-31T22:16:00+00:00\",\"dateModified\":\"2026-08-01T17:59:46+00:00\",\"description\":\"As agentic and long&#x2d;context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1).\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/llm-optimize-deploy.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/llm-optimize-deploy.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/31\\\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference - Future News 24","description":"As agentic and long&#x2d;context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1).","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/","og_locale":"en_US","og_type":"article","og_title":"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference - Future News 24","og_description":"As agentic and long&#x2d;context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1).","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/","og_site_name":"Future News 24","article_published_time":"2026-07-31T22:16:00+00:00","article_modified_time":"2026-08-01T17:59:46+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/llm-optimize-deploy.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/llm-optimize-deploy.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"12 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference","datePublished":"2026-07-31T22:16:00+00:00","dateModified":"2026-08-01T17:59:46+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/"},"wordCount":2347,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/llm-optimize-deploy.webp","keywords":["Attention","CoDesigning","Fast","inference","interactive","LongContext","Model"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/","name":"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/llm-optimize-deploy.webp","datePublished":"2026-07-31T22:16:00+00:00","dateModified":"2026-08-01T17:59:46+00:00","description":"As agentic and long&#x2d;context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1).","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/llm-optimize-deploy.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/llm-optimize-deploy.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/31\/co-designing-ai-model-attention-for-fast-interactive-long-context-inference\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Co-Designing AI Mannequin Consideration for Quick, Interactive Lengthy-Context Inference"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3146","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=3146"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3146\/revisions"}],"predecessor-version":[{"id":3147,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3146\/revisions\/3147"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/3148"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=3146"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=3146"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=3146"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}