{"id":2213,"date":"2026-07-10T00:00:00","date_gmt":"2026-07-10T00:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/"},"modified":"2026-07-12T01:59:31","modified_gmt":"2026-07-12T01:59:31","slug":"torch-attention-profile","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/","title":{"rendered":"Profiling in PyTorch (Half 3): Consideration is all you profile"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/profile-3-thumbnail.png\" alt=\"Thumbnail of the blog post\"\/><\/p>\n<p>The collection &#8220;Profiling in PyTorch&#8221; is supposed to make you comfy studying profiler traces and tables. In Half 1 we profiled primary math operations like addition and multiplication. We noticed how the profiler desk uncovers hotspots, and the way the profiler hint reveals the order through which an algorithm runs over time.<\/p>\n<p>In Half 2 we wrapped that addition and multiplication right into a torch linear layer. We then stacked a number of linear layers on prime of one another (a multilayer perceptron) and profiled that. Alongside the way in which we additionally profiled fused and hand-tuned kernels.<\/p>\n<p>From the attitude of the Transformer structure, the subsequent logical step for us to profile is one more basic algorithm, consideration. Whereas being notorious for its quadratic-time complexity, many intelligent methods exist to mitigate that challenge and make it quick. Our objective right here is to not cowl each trick intimately. As an alternative, we wish to see how each appears completely different beneath the profiler.<\/p>\n<blockquote class=\"note\">\n<p>The scripts for this weblog publish stay right here: 04_a_naive_attention.py, 04_b_inplace_ops_attention.py, 04_c_sdpa_attention.py, and 04_d_kernels_attention.py. Like earlier than, it helps to open them in a separate tab and stroll by means of the code as you learn. We use an NVIDIA A100-SXM4-80GB GPU to run the scripts. It&#8217;s very easy to arrange a GPU on the Hugging Face infrastructure and experiment with the scripts utilizing Dev Mode with Areas. One may additionally run the scripts with the Hugging Face Jobs pipeline.<\/p>\n<\/blockquote>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tNaive consideration<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Consideration works with Queries (q), Keys (ok), and Values (v). The interplay between them might be written as a brief sequence of steps:<\/p>\n<p>Construct the eye scores scores: matmul(q, ok.T)<br \/>\nScale the scores: scores * scale<br \/>\nApply a causal masks to the scores: scores.masked_fill(masks, &#8220;-inf&#8221;)<br \/>\nNormalize the scores with softmax to get the eye weights attn: softmax(scores)<br \/>\nReweight the values with these weights: matmul(attn, v)<\/p>\n<p>So consideration can be a assortment of primitive operations. A few of them we already know (the matmuls), and the remainder are simple to identify. Let&#8217;s write a naive consideration module in PyTorch and profile it.<\/p>\n<p><span class=\"hljs-keyword\">class<\/span> <span class=\"hljs-title class_\">NaiveCausalAttention<\/span>(nn.Module):<br \/>\n    <span class=\"hljs-keyword\">def<\/span> <span class=\"hljs-title function_\">__init__<\/span>(<span class=\"hljs-params\">self, head_dim<\/span>):<br \/>\n        <span class=\"hljs-built_in\">tremendous<\/span>().__init__()<br \/>\n        self.scale = <span class=\"hljs-number\">1.0<\/span> \/ math.sqrt(head_dim)<\/p>\n<p>    <span class=\"hljs-keyword\">def<\/span> <span class=\"hljs-title function_\">ahead<\/span>(<span class=\"hljs-params\">self, q, ok, v, masks<\/span>):<br \/>\n        scores = torch.matmul(q, ok.transpose(-<span class=\"hljs-number\">2<\/span>, &#8211;<span class=\"hljs-number\">1<\/span>))<br \/>\n        scores = scores * self.scale<br \/>\n        scores = scores.masked_fill(masks, <span class=\"hljs-built_in\">float<\/span>(<span class=\"hljs-string\">&#8220;-inf&#8221;<\/span>))<br \/>\n        attn = torch.softmax(scores, dim=-<span class=\"hljs-number\">1<\/span>)<br \/>\n        out = torch.matmul(attn, v)<br \/>\n        <span class=\"hljs-keyword\">return<\/span> out<\/p>\n<p>Earlier than opening the hint, let&#8217;s do our traditional train and guess what we should always see. Tracing the ahead of this module, we anticipate:<\/p>\n<p>a matmul kernel (q . ok.T)<br \/>\na mul kernel (the scaling)<br \/>\nan operation for the masking<br \/>\na softmax kernel<br \/>\na matmul kernel (atten . v)<\/p>\n<p>uv run 04_a_naive_attention.py<br \/>\nuvx trace-util -f traces\/ -b \/traces<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/cpu-profile-naive.png\" alt=\"CPU lane of the naive attention profiler trace, with the `attn_fwd` block expanded to show its matmul, mul, masked_fill and softmax operations\"\/><\/p>\n<p>Determine 1: The CPU lane of the profile hint for naive consideration highlighting the discrete operations<\/p>\n<\/div>\n<p>Determine 1 reveals the CPU lane of the profile (the GPU lane is folded so it doesn&#8217;t overwhelm us). Inside attn_fwd (our annotated ahead name) we will see precisely the operations we guessed. The matmul is an previous good friend by now, and the brand new operations are simple to identify:<\/p>\n<p>mul: the scaling<br \/>\nmasked_fill: the causal masking<br \/>\nsoftmax: the softmax kernel<\/p>\n<p>Now let&#8217;s unfold the GPU lane and see which kernels have been truly launched.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/gpu-profile-naive.png\" alt=\"Profiler trace of naive attention showing the CPU lane above the GPU lane, with each `attn_fwd` step mapping to a cluster of GPU kernels\"\/><\/p>\n<p>Determine 2: GPU and CPU lanes of the profile hint for naive consideration highlighting a set of kernels corresponding to at least one profiler step.<\/p>\n<\/div>\n<p>Determine 2 reveals the GPU lane subsequent to the CPU lane. Let&#8217;s zoom right into a single attn_fwd block on the GPU lane to have a look at the kernels one after the other.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/each-kernels-naive.png\" alt=\"Zoomed-in GPU lane of naive attention showing the individual kernels for one step: two matmuls, a mul, a memory copy, a masking kernel and a softmax\"\/><\/p>\n<p>Determine 3: Zoomed in GPU lane of the profiler hint for naive consideration implementation.<\/p>\n<\/div>\n<p>Determine 3 lets us learn off the person kernels for one profiler step:<\/p>\n<p>matmul (question and key)<br \/>\nmul (scaling)<br \/>\nreminiscence copy \ud83e\udd14<br \/>\ncausal masking<br \/>\nsoftmax (produces the eye weights)<br \/>\nmatmul (consideration weights and values)<\/p>\n<p>5 of those are anticipated. The reminiscence copy is the odd one out, so the place does this come from? The clue is that PyTorch has in-place operations. While you function on a tensor the unusual (out-of-place) manner, PyTorch usually makes a replica, applies the requested operation to it, and returns the copy. Following the sequence of operations, the offender right here is our masked_fill.<\/p>\n<p>What if we changed this with an in-place operation?<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tNaive consideration with inplace causal masking<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>All we alter is masked_fill to masked_fill_ (word the trailing underscore, PyTorch&#8217;s conference for in-place operations), and we run the identical script.<\/p>\n<p>def ahead(self, q, ok, v, masks):<br \/>\n    # q, ok, v: [batch, heads, seq, head_dim]<br \/>\n    scores = torch.matmul(q, ok.transpose(-2, -1))  # [batch, heads, seq, seq]<br \/>\n    scores = torch.mul(scores, self.scale)<br \/>\n<span class=\"hljs-deletion\">&#8211;    scores = scores.masked_fill(masks, float(&#8220;-inf&#8221;))<\/span><br \/>\n<span class=\"hljs-addition\">+    scores.masked_fill_(masks, float(&#8220;-inf&#8221;))<\/span><br \/>\n    attn = torch.softmax(scores, dim=-1)<br \/>\n    out = torch.matmul(attn, v)  # [batch, heads, seq, head_dim]<br \/>\n    return out<\/p>\n<p>Let us take a look at the hint and see if one thing modified.<\/p>\n<p>uv run 04_b_inplace_ops_attention.py<br \/>\nuvx trace-util -f traces\/ -b \/traces<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Kind<br \/>\nCPU stream<\/p>\n<p>Determine 4: Naive masking<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/cpu-profile-naive.png\" alt=\"CPU lane of naive attention with out-of-place `masked_fill`, showing several dispatch ops for the masking step\"\/><\/p>\n<p>Determine 5: In place masking<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/cpu-profile-inplace.png\" alt=\"CPU lane of naive attention with in-place `masked_fill_`, showing fewer dispatch ops for the masking step\"\/><\/p>\n<\/div>\n<p>The in-place model (Determine 5) wraps far fewer CPU ops contained in the masking step than the out-of-place model (Determine 4). That is an encouraging sign. Let&#8217;s unfold the GPU lane to verify what occurred there.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Kind<br \/>\nGPU stream<\/p>\n<p>Determine 6: Naive masking<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/each-kernels-naive.png\" alt=\"GPU kernels for naive attention including a separate Memcpy kernel before the masking\"\/><\/p>\n<p>Determine 7: In place masking<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/each-kernels-inpace.png\" alt=\"GPU kernels for naive attention with in-place masking, with the Memcpy kernel gone\"\/><\/p>\n<\/div>\n<p>On the GPU lane the Memcpy kernel is gone for good (Figures 6 and seven). With a one line change we shaved a complete kernel off every ahead move. This may occasionally not seem like a lot by itself, however keep in mind it is a single consideration operation. Within the context of a transformer based mostly giant mannequin (LLMs, Diffusion fashions, and so forth.), it repeats as soon as per layer, and there are a lot of layers, so the saving provides up shortly (and if it earns you a increase, sharing at the least 10% with us feels solely truthful).<\/p>\n<blockquote class=\"note\">\n<p>Out-of-place is PyTorch&#8217;s default for a cause. To compute gradients, autograd has to recollect the tensor values it noticed on the ahead move, as a result of many backward formulation reuse them. An in-place operation overwrites these values in reminiscence, so the backward move would learn the mistaken numbers. Because of the truth that we run ahead beneath torch.no_grad, in-place is secure for us, with no backward move and nothing to deprave. Additionally it is noteworthy that in-place operations don&#8217;t solely save time (like we see in our case) but in addition reminiscence (as a consequence of no further copy) which is nice for big tensors like logits!<\/p>\n<\/blockquote>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tScaled Dot Product Consideration<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>We simply constructed consideration from primitives, and even shaved off a Memcpy. The excellent news is that the PyTorch group has finished all of this for us, and packaged the entire pipeline right into a single operate:<\/p>\n<p><span class=\"hljs-keyword\">from<\/span> torch.nn <span class=\"hljs-keyword\">import<\/span> useful <span class=\"hljs-keyword\">as<\/span> F<\/p>\n<p>F.scaled_dot_product_attention(q, ok, v, is_causal=<span class=\"hljs-literal\">True<\/span>)<\/p>\n<p>This one line replaces our hand written module, and is_causal=True even saves us from constructing the masks by hand. It&#8217;s price pausing to understand how a lot this one name hides. And it hides extra than simply code traces. Scaled Dot Product Consideration (SDPA) doesn&#8217;t have a single implementation. Below the hood it dispatches to one of many a number of backends and picks the quickest one which helps our inputs (dtype, head dimension, masks, {hardware}, and so forth.).<\/p>\n<p>The official SDPA tutorial walks us by means of this choice, and the backends themselves are listed within the torch.nn.consideration.SDPBackend enum:<\/p>\n<p><span class=\"hljs-keyword\">from<\/span> torch.nn.consideration <span class=\"hljs-keyword\">import<\/span> SDPBackend<\/p>\n<p>BACKENDS = {<br \/>\n    <span class=\"hljs-string\">&#8220;math&#8221;<\/span>: SDPBackend.MATH,<br \/>\n    <span class=\"hljs-string\">&#8220;flash&#8221;<\/span>: SDPBackend.FLASH_ATTENTION,<br \/>\n    <span class=\"hljs-string\">&#8220;environment friendly&#8221;<\/span>: SDPBackend.EFFICIENT_ATTENTION,<br \/>\n    <span class=\"hljs-string\">&#8220;cudnn&#8221;<\/span>: SDPBackend.CUDNN_ATTENTION,<br \/>\n}<\/p>\n<p>Usually SDPA chooses for us, however we will pin a particular backend with the torch.nn.consideration.sdpa_kernel context supervisor. That is what we do in our scripts. This lets us profile every backend by itself and skim how in a different way they present up within the hint. Let&#8217;s go one by one.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tMath backend<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>uv run 04_c_sdpa_attention.py &#8211;backend math<br \/>\nuvx trace-util -f traces\/ -b \/traces<\/p>\n<p>Earlier than we open something, let&#8217;s guess. We&#8217;ve changed hand written consideration (matmul, mul, masks, softmax, matmul) with a single one liner, so we should always anticipate the hint to get easier and sooner. Fewer kernels, much less CPU dispatch, perhaps even a fused kernel. Let&#8217;s verify the profiler desk first.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Metric<br \/>\nThe place to look?<br \/>\nNaive in-place<br \/>\nSDPA math<\/p>\n<p>*_fwd CUDA time avg<br \/>\nThe &#8220;CUDA time avg&#8221; column for the *_fwd op<br \/>\n1.955 ms<br \/>\n7.239 ms<\/p>\n<p>Self CUDA time complete<br \/>\nOn the backside of the profiler desk<br \/>\n7.194 ms<br \/>\n27.279 ms<\/p>\n<\/div>\n<p>That is our first shock, the one liner is 3.7x slower.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Profiler Hint<\/p>\n<p>Determine 8: Profiler hint of naive in-place consideration exhibiting 5 GPU kernel launches for one ahead<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/inplace-kernel-launches.png\" alt=\"GPU lane of naive in-place attention with five kernel launches for one forward pass\"\/><\/p>\n<p>Determine 9: Profiler hint of the SDPA math backend exhibiting 20 GPU kernel launches for a single consideration ahead<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/math-kernel-launches.png\" alt=\"GPU lane of the SDPA math backend with twenty kernel launches for a single attention forward pass\"\/><\/p>\n<\/div>\n<p>Opening the hint (Determine 9) reveals why the alarm bells ring, the mathematics backend launches 20 GPU kernels per ahead as a substitute of the 5 launched with our naive consideration implementation (Determine 8). That is the other of what we guessed. Let&#8217;s work out why this occurs.<\/p>\n<h4 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tTensor cores left vacant<br \/>\n\t<\/span><br \/>\n<\/h4>\n<p>In Half 2 we discovered to learn a kernel title like a fingerprint. Let&#8217;s use that behavior right here:<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Run<br \/>\nmatmul kernel<\/p>\n<p>Determine 10: Naive consideration<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/tensor-core-kernels.png\" alt=\"Matmul kernel name for naive attention in Perfetto, carrying the s16816 bfloat16 Tensor-core GEMM signature\"\/><\/p>\n<p>Determine 11: SDPA with math backend<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/cuda-core-kernels.png\" alt=\"Matmul kernel name for the SDPA math backend, carrying the sgemm FP32 CUDA-core signature\"\/><\/p>\n<\/div>\n<p>The A100s we used to seize these traces ship with Tensor Cores, specialised {hardware} for accelerated matmuls that&#8217;s identified to be far sooner than the unusual CUDA cores. To see why that issues right here, it helps to know what lives inside a GPU. A Streaming Multiprocessor (SM) is the compute unit of a GPU, and every SM has two sorts of arithmetic models, the CUDA cores and the Tensor Cores. CUDA cores are common objective and course of a handful of components at a time, whereas Tensor Cores multiply and accumulate a complete small matrix tile in a single instruction. So the query is easy, &#8220;Is every backend truly utilizing the quick path?&#8221;<\/p>\n<p>The kernel names reply it. The s16816 within the naive kernel (Determine 10) is the signature of a bfloat16 Tensor Core matmul (the 16x8x16 Tensor Core instruction), so the naive model is on the quick path. sgemm (Determine 11) is the basic single precision (FP32) matmul that runs on the unusual CUDA cores. In different phrases, the mathematics backend by no means touches the Tensor Cores in any respect: to commerce pace for numerical accuracy it upcasts tensors to FP32 (doubling the info moved, even when the inputs are in bf16) and falls again to the slower CUDA cores.<\/p>\n<h4 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tCausal masks constructed<br \/>\n\t<\/span><br \/>\n<\/h4>\n<p>Within the naive model we constructed the causal masks as soon as and reused it. Right here we handed is_causal=True and the mathematics backend materialized one for us, on each single name. You may watch it occur on the CPU lane:<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/mask-math.png\" alt=\"CPU lane of the SDPA math backend showing the ops that rebuild the causal mask: aten::ones, aten::tril, aten::scalar_tensor, aten::fill_ and aten::where\"\/><\/p>\n<p>Determine 12: CPU lane exhibiting the ops for masking<\/p>\n<\/div>\n<p>Here&#8217;s what we see in Determine 12<\/p>\n<p>aten::ones -&gt; aten::tril            construct a [<span class=\"hljs-built_in\">seq<\/span>, <span class=\"hljs-built_in\">seq<\/span>] lower-triangular matrix<br \/>\naten::scalar_tensor -&gt; aten::fill_  make the -inf fill worth<br \/>\naten::<span class=\"hljs-built_in\">the place<\/span>                         flip it into an additive bias (0 or -inf)<\/p>\n<p>On the GPU this reveals up as a triu_tril_kernel, a number of the place kernels, and an add_. The comfort flag that permit us cease desirous about the masks didn&#8217;t take away the work, it simply moved it one layer down, the place the masks is rebuilt from scratch each ahead.<\/p>\n<h4 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tThe secure softmax<br \/>\n\t<\/span><br \/>\n<\/h4>\n<p>Our hand written model known as plain aten::softmax. The mathematics backend calls aten::_safe_softmax, and the distinction is once more seen as further kernels (Determine 13):<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/safe-softmax-extra-kernels.png\" alt=\"GPU lane of the SDPA math backend showing the extra kernels that aten::_safe_softmax launches compared to a plain softmax\"\/><\/p>\n<p>Determine 13: Secure softmax highlighting the additional kernels in comparison with generic softmax<\/p>\n<\/div>\n<p>A row that&#8217;s totally masked (each entry -inf) would make an unusual softmax compute exp(-inf)\/sum(exp(-inf)) = 0\/0 = NaN. _safe_softmax guards towards precisely that. Our naive kernel by no means bothered, and would have quietly produced NaNs in that nook case.<\/p>\n<h4 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tSo what&#8217;s the math backend for?<br \/>\n\t<\/span><br \/>\n<\/h4>\n<p>Put collectively, the mathematics backend is the reference implementation. It&#8217;s a easy, dtype-safe, NaN-safe decomposition of consideration into primitive ATen ops. It&#8217;s basically the naive consideration we wrote by hand, however extra cautious. That carefulness is precisely what makes it extraordinarily sluggish.<\/p>\n<p>Its job is to not be quick, however to at all times work. This makes it the right baseline. Each backend we profile subsequent (flash, environment friendly, cudnn) is making an attempt to break down the 20 GPU kernels into basically one fused kernel that stays in bf16 and by no means materializes the intermediate matrices in any respect.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tEnvironment friendly backend<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>uv run 04_c_sdpa_attention.py &#8211;backend environment friendly<br \/>\nuvx trace-util -f traces -b \/traces<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/efficient-backend.png\" alt=\"Profiler trace of the SDPA efficient backend showing a single fused fmha_cutlassF attention kernel per forward\"\/><\/p>\n<p>Determine 14: The profiler hint for sdpa with environment friendly backend<\/p>\n<\/div>\n<p>The place the mathematics backend launched 20 kernels throughout one profiler step, the environment friendly backend launches just one fmha_cutlassF_bf16_aligned_64x64_rf_sm80 (as seen in Determine 14).<\/p>\n<p>Let&#8217;s decode the title of the kernel:<\/p>\n<p>fmha (fused multi-head consideration): All of the primitive ops in consideration is &#8220;fused&#8221; in a single op now.<br \/>\ncutlassF: constructed on CUTLASS (NVIDIA&#8217;s open-source templates for tensor-core GEMMs), F for ahead.<br \/>\nbf16_aligned: runs in bfloat16 (no FP32 upcast, not like math).<br \/>\n64&#215;64: the tile measurement.<br \/>\nrf (register file): the working set is saved in registers, the quickest reminiscence on the chip.<br \/>\nsm80: compiled for Ampere (the A100&#8217;s compute functionality 8.0).<\/p>\n<p>That is the reminiscence environment friendly consideration kernel that grew out of Meta&#8217;s xformers library and was upstreamed into PyTorch. When individuals say &#8220;the xformers backend,&#8221; this fmha_cutlassF kernel is what they imply.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tFlash backend<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>uv run 04_c_sdpa_attention.py &#8211;backend flash<br \/>\nuvx trace-util -f traces -b \/traces<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/flash-backend.png\" alt=\"Profiler trace of the SDPA flash backend\"\/><\/p>\n<p>Determine 15: The flash backend hint, one fused pytorch_flash kernel per ahead<\/p>\n<\/div>\n<p>The void pytorch_flash kernel (Determine 15) is FlashAttention-2 (Tri Dao&#8217;s implementation), vendored into PyTorch.<\/p>\n<p>Earlier than we learn the hint any additional, it&#8217;s price answering the query you ought to be asking by now: why is there a complete backend named &#8220;flash&#8221;, and why does it matter a lot?<\/p>\n<h4 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tWhy flash consideration exists?<br \/>\n\t<\/span><br \/>\n<\/h4>\n<p>Let&#8217;s return to the mathematics backend for a second. Its actual drawback was not the depend of 20 kernels, it was what these kernels handed to one another.<\/p>\n<p>Step 1 builds the complete rating matrix attn = q . ok.T, which is [seq, seq] per head. For a sequence size of 4096 that&#8217;s 4096 x 4096 \u2248 16 million numbers for a single head. That matrix is written out to the HBM (the GPU&#8217;s major reminiscence), if there may be even sufficient area to take action. Then, it&#8217;s learn again to be scaled, written once more for the masks, learn once more for the softmax, and so forth. Consideration&#8217;s price is dominated by this forwards and backwards site visitors to HBM, not by the matmuls themselves.<\/p>\n<p>FlashAttention assaults precisely this. As an alternative of computing the entire s matrix and solely then decreasing it, it walks over ok and v in tiles, retains a operating softmax because it goes (the &#8220;on-line softmax&#8221; trick), and accumulates the output one tile at a time. The complete [seq, seq] rating matrix is rarely written to HBM, it solely ever lives on-chip. That is the only concept that lets the whole consideration pipeline collapse into one fused kernel that stays in bf16 on the Tensor cores.<\/p>\n<h4 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tWhy flash appears &#8220;mistaken&#8221; beneath the profiler<br \/>\n\t<\/span><br \/>\n<\/h4>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/flash-occupancy.png\" alt=\"Perfetto footprint of the flash kernel reporting an estimated achieved occupancy of 13%\"\/><\/p>\n<p>Determine 16: Estimated occupancy of flash kernel is seen to be 13%<\/p>\n<\/div>\n<p>Right here is the place flash surprises individuals who learn profiler footprints. It&#8217;s the quickest backend, but the profiler studies it with very low occupancy (proven in Determine 16). To see why that&#8217;s high-quality, we&#8217;d like three fast definitions.<\/p>\n<p>A GPU kernel is actually a collection of directions executed by many small execution models. These particular person execution models (threads) care for loading variables, including them collectively, storing them again, and so forth. For every kernel, we launch many, many threads, and to maintain observe of them, we group them by blocks.<\/p>\n<p>Blocks are scheduled onto Streaming Multiprocessors (SMs), the principle compute models of a GPU. A block lives fully on one SM, and an SM can host a number of blocks without delay if it has sufficient sources. These sources embody registers, shared reminiscence, most resident threads, and most resident warps. So after we say a kernel has low occupancy, we imply every SM has fewer resident warps than it may theoretically assist.<\/p>\n<blockquote class=\"tip\">\n<p>If you wish to know extra about threads, blocks, grids, and so forth. right here is a superb useful resource.<\/p>\n<\/blockquote>\n<p>In case you click on the flash kernel within the hint, its footprint tells the story (Determine 17).<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/flash-reg-count.png\" alt=\"Resource footprint of the pytorch_flash kernel in Perfetto, showing a high per-thread register count and large shared memory usage per block\"\/><\/p>\n<p>Determine 17: The flash kernel footprint, heavy on registers and shared reminiscence per block.<\/p>\n<\/div>\n<p>Flash makes use of a variety of per-thread registers and a considerable amount of shared reminiscence per block. For instance, if a block has 128 threads and every thread makes use of 255 registers, that block wants 128 \u00d7 255 = 32,640 registers. On an Ampere SM with 65,536 registers, solely two such blocks match without delay. Every 128-thread block has 128 \/ 32 = 4 warps, so two blocks give solely 8 resident warps. Towards a most of 64 resident warps, that&#8217;s roughly 13% occupancy. Flash has low occupancy not as a result of it&#8217;s poorly optimized, however as a result of every block is intentionally very &#8220;heavy&#8221; in on-chip useful resource utilization.<\/p>\n<p>And that&#8217;s the complete level. Excessive occupancy helps disguise latency by holding many warps able to run, however it doesn&#8217;t make the work itself environment friendly. Flash spends these registers and that shared reminiscence on objective, to maintain consideration tiles on-chip, reuse knowledge aggressively, and keep away from ever materializing the complete consideration matrix in world reminiscence.<\/p>\n<h3 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tcuDNN backend<br \/>\n\t<\/span><br \/>\n<\/h3>\n<p>uv run 04_c_sdpa_attention.py &#8211;backend cudnn<br \/>\nuvx trace-util -f traces -b \/traces<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/cudnn-backend.png\" alt=\"Profiler trace of the SDPA cuDNN backend showing a single cudnn_generated attention kernel per forward\"\/><\/p>\n<p>Determine 18: The cuDNN backend hint, a single generated consideration kernel per ahead.<\/p>\n<\/div>\n<p>By now the sample is acquainted. Like flash and environment friendly, cuDNN offers us one fused, flash-style kernel per ahead (Determine 18). So the pure query is: if flash already fuses consideration, why does PyTorch ship one more flash backend? The reply is who writes the kernel and the way it&#8217;s constructed, and that distinction is what makes the hint look completely different.<\/p>\n<h4 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tHow is cuDNN kernel completely different<br \/>\n\t<\/span><br \/>\n<\/h4>\n<p>Flash and environment friendly are fastened, pre-compiled kernels vendored into PyTorch. You get the identical binary each time. cuDNN is NVIDIA&#8217;s personal deep studying library, and its consideration kernel is generated and tuned for the precise drawback at hand. It&#8217;s nearer in spirit to torch.compile&#8217;s codegen than to a hard and fast cuBLAS binary. You may learn that straight off the (very lengthy) kernel title:<\/p>\n<p>cudnn_generated_fort_native_sdpa_sm80_flash_fprop_wmma_f16_knob_6_128x64x64_4x1x1_cga1x1x1_kernel0_0<\/p>\n<p>cudnn_generated: not a pre-shipped binary, it was generated by cuDNN.<br \/>\nflash_fprop: a flash consideration type ahead move. So the algorithm is similar household because the flash backend.<br \/>\nwmma_f16: it makes use of the warp-level matrix multiply-accumulate (WMMA) API, the Tensor-core path on the 16-bit float pipeline.<br \/>\nknob_6: cuDNN picks from a set of pre-tuned configurations (&#8220;knobs&#8221;). Completely different shapes choose completely different knobs, very like cuBLAS choosing a tile variant.<br \/>\n128x64x64: the tile dimensions it selected.<\/p>\n<p>That one truth, generated per drawback, explains every part else that appears uncommon within the hint.<\/p>\n<p>No transposes: The CPU lane goes from _cudnn_attention_forward straight to a few aten::empty allocations after which the kernel, with zero aten::transpose (Figures 19, 20 and 21). Flash and environment friendly every insert 4 (metadata) transposes to reshape the tensors whereas cuDNN consumes the native [B, H, S, D] format straight as a result of its generator emits a kernel for that format.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Variant<br \/>\nHint<\/p>\n<p>Determine 19: Flash<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/flash-transpose.png\" alt=\"CPU lane of the flash backend showing four aten::transpose ops before the fused attention kernel\"\/><\/p>\n<p>Determine 20: Environment friendly<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/efficient-trasnpose.png\" alt=\"CPU lane of the efficient backend showing four aten::transpose ops before the fused attention kernel\"\/><\/p>\n<p>Determine 21: cuDNN<br \/>\n<img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/cudnn-backend.png\" alt=\"CPU lane of the cuDNN backend going straight to aten::empty allocations and the kernel, with no transpose ops\"\/><\/p>\n<\/div>\n<p>It launches by means of cuLaunchKernelEx, not cudaLaunchKernel: Each different kernel on this complete collection went by means of the runtime API cudaLaunchKernel. cuDNN makes use of the driver-level prolonged launch, which carries launch attributes (Determine 22).<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/cudnn-launch.png\" alt=\"CPU lane of the cuDNN backend showing the cuLaunchKernelEx driver-level launch instead of cudaLaunchKernel\"\/><\/p>\n<p>Determine 22: CPU lane of the cuDNN backend exhibiting the cuLaunchKernelEx driver-level launch as a substitute of cudaLaunchKernel<\/p>\n<\/div>\n<p>The profiler studies 0% achieved occupancy: Don&#8217;t take that at face worth, it&#8217;s a measurement hole, not a stalled GPU. CUPTI (the profiling backend) can&#8217;t attribute occupancy to a driver-API (cuLaunchKernelEx) launch the way in which it does for cudaLaunchKernel, so the sphere reads 0. The footprint fills within the fact (Determine 23): 240 registers \u00d7 256 threads = 61,440 registers per block towards the SM&#8217;s 65,536, so just one block suits per SM (8 warps \u2248 12.5%), proper consistent with flash.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p><img decoding=\"async\" src=\"https:\/\/huggingface.co\/datasets\/huggingface\/documentation-images\/resolve\/main\/blog\/torch-attention-profile\/cudnn-footprint.png\" alt=\"Perfetto footprint of the cuDNN kernel reporting 0% achieved occupancy, with 240 registers per thread and 256 threads per block\"\/><\/p>\n<p>Determine 23: cuDNN kernel reporting 0% achieved occupancy, with 240 registers per thread and 256 threads per block<\/p>\n<\/div>\n<h4 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tThe price moved to the CPU<br \/>\n\t<\/span><br \/>\n<\/h4>\n<p>The &#8220;no transposes&#8221; story tempts us to anticipate cuDNN to be the leanest backend on the CPU. It&#8217;s the reverse.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>backend<br \/>\nCUDA avg time<br \/>\nCPU avg time<\/p>\n<p>environment friendly<br \/>\n277.9 \u00b5s<br \/>\n117 \u00b5s<\/p>\n<p>flash<br \/>\n146.8 \u00b5s<br \/>\n138 \u00b5s<\/p>\n<p>cudnn<br \/>\n186.3 \u00b5s<br \/>\n214 \u00b5s<\/p>\n<\/div>\n<p>Even with zero transpose ops, cuDNN spends about 214 \u00b5s per ahead on the CPU, greater than flash (138) or environment friendly (117). Nearly all of it sits in aten::scaled_dot_product_attention self time (26% of the entire run) and _cudnn_attention_forward. That&#8217;s cuDNN&#8217;s runtime engine choosing and making ready the plan (the &#8220;knob&#8221; search) on each name.<\/p>\n<p>Fewer seen ATen ops didn&#8217;t imply much less CPU work, it moved the work into the library, the place the profiler can solely present it as one fats, opaque bar. When a hint immediately will get cleaner, the work has not at all times disappeared, generally it has simply moved someplace the profiler can&#8217;t break down.<\/p>\n<p>On the GPU, cuDNN (186.3 \u00b5s) lands between environment friendly and flash. On this very flash-friendly form, hand-written FlashAttention-2 edges it out. cuDNN usually wins on different shapes (bigger head dimensions, completely different sequence lengths) exactly as a result of its generator retunes per drawback, however that retuning can also be what you simply paid for on the CPU.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tAll the things we lined, at a look<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Earlier than we wrap up, here&#8217;s a single desk to evaluate each consideration variant we profiled and the one lesson every hint taught us.<\/p>\n<div class=\"max-w-full overflow-auto\">\n<p>Variant<br \/>\nWhat we modified<br \/>\nKernels \/ ahead<br \/>\nWhat the hint revealed<\/p>\n<p>Naive consideration<br \/>\nConsideration constructed by hand from primitives (matmul, mul, masks, softmax, matmul)<br \/>\n6<br \/>\nA hidden Memcpy from the out-of-place masked_fill.<\/p>\n<p>Naive in-place<br \/>\nmasked_fill \u2192 masked_fill_<br \/>\n5<br \/>\nOne line drops the Memcpy kernel fully.<\/p>\n<p>SDPA math<br \/>\nF.scaled_dot_product_attention pinned to the mathematics backend<br \/>\n20<br \/>\nThe reference: FP32 on CUDA cores, masks rebuilt each name, _safe_softmax. Appropriate however ~3.7x slower.<\/p>\n<p>SDPA environment friendly<br \/>\nEnvironment friendly (xformers) backend<br \/>\n1<br \/>\nOne fused fmha_cutlassF kernel, stays in bf16 on Tensor cores.<\/p>\n<p>SDPA flash<br \/>\nFlash backend<br \/>\n1<br \/>\nOne fused pytorch_flash kernel (FlashAttention-2). Quickest, regardless of &#8220;wrong-looking&#8221; 13% occupancy.<\/p>\n<p>SDPA cuDNN<br \/>\ncuDNN backend<br \/>\n1<br \/>\nA per-problem generated kernel: no transposes, cuLaunchKernelEx, however the fee moved to a fats CPU bar.<\/p>\n<\/div>\n<h2 class=\"relative group flex items-baseline\">\n<p>\t<span><br \/>\n\t\tConcluding the collection<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>In case you take away just one factor from the entire collection, let it&#8217;s the behavior we repeated earlier than each single hint which is to guess first, then look.<\/p>\n<p>State out loud what you anticipate the hint to comprise, open it, and deal with any mismatch as probably the most attention-grabbing factor on the display. Each actual perception in these three posts, the hidden Memcpy, the addmm epilogue, the 20 kernel math backend, flash&#8217;s &#8220;wrong-looking&#8221; occupancy, cuDNN&#8217;s fats CPU bar, got here from a guess that didn&#8217;t match the hint.<\/p>\n<p>Profiling shouldn&#8217;t be a separate, intimidating talent reserved for GPU consultants. It&#8217;s simply the self-discipline of trying carefully and asking &#8220;wait, why is that taking place?&#8221; till the reply clicks. You now have the vocabulary and the reflexes to try this by yourself fashions. Open a hint, type a guess, and go discover the mismatch.<\/p>\n<p>Thanks for studying the Profiling in PyTorch collection. Now go profile one thing. \ud83e\udd17<\/p>\n<p>Due to Noe Flandre for his or her critiques on the early draft of the publish!<\/p>\n<blockquote class=\"note\">\n<p>The weblog publish was polished utilizing an LLM. This by no means signifies that we have now let an agent run within the background and let it generate the weblog. A few of us within the group are non-english audio system and suppose LLMs (that are principally skilled within the English Language) can rectify foolish grammar errors or rephrase sentences that sound much less intimidating and cleaner. Hope this helps with the concept of &#8220;why ought to I learn, if this was LLM generated&#8221;. \ud83e\udd17<\/p>\n<\/blockquote>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/huggingface.co\/blog\/torch-attention-profile\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The collection &#8220;Profiling in PyTorch&#8221; is supposed to make you comfy studying profiler traces and tables. In Half 1 we profiled primary math operations like addition and multiplication. We noticed how the profiler desk uncovers hotspots, and the way the profiler hint reveals the order through which an algorithm runs over time. In Half 2 [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":2215,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/huggingface.co\/blog\/assets\/torch-attention-profile\/thumbnail.png","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[5],"tags":[1050,1227,2733,1225,1226],"class_list":["post-2213","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-developer-ai-open-source-ecosystem","tag-attention","tag-part","tag-profile","tag-profiling","tag-pytorch"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Profiling in PyTorch (Half 3): Consideration is all you profile - Future News 24<\/title>\n<meta name=\"description\" content=\"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Profiling in PyTorch (Half 3): Consideration is all you profile - Future News 24\" \/>\n<meta property=\"og:description\" content=\"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-10T00:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-12T01:59:31+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/huggingface.co\/blog\/assets\/torch-attention-profile\/thumbnail.png\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/huggingface.co\/blog\/assets\/torch-attention-profile\/thumbnail.png\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"21 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Profiling in PyTorch (Half 3): Consideration is all you profile\",\"datePublished\":\"2026-07-10T00:00:00+00:00\",\"dateModified\":\"2026-07-12T01:59:31+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/\"},\"wordCount\":4189,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/huggingface.co\\\/blog\\\/assets\\\/torch-attention-profile\\\/thumbnail.png\",\"keywords\":[\"Attention\",\"Part\",\"profile\",\"Profiling\",\"PyTorch\"],\"articleSection\":[\"Developer AI &amp; Open-Source Ecosystem\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/\",\"name\":\"Profiling in PyTorch (Half 3): Consideration is all you profile - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/huggingface.co\\\/blog\\\/assets\\\/torch-attention-profile\\\/thumbnail.png\",\"datePublished\":\"2026-07-10T00:00:00+00:00\",\"dateModified\":\"2026-07-12T01:59:31+00:00\",\"description\":\"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/#primaryimage\",\"url\":\"https:\\\/\\\/huggingface.co\\\/blog\\\/assets\\\/torch-attention-profile\\\/thumbnail.png\",\"contentUrl\":\"https:\\\/\\\/huggingface.co\\\/blog\\\/assets\\\/torch-attention-profile\\\/thumbnail.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/10\\\/torch-attention-profile\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Profiling in PyTorch (Half 3): Consideration is all you profile\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Profiling in PyTorch (Half 3): Consideration is all you profile - Future News 24","description":"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/","og_locale":"en_US","og_type":"article","og_title":"Profiling in PyTorch (Half 3): Consideration is all you profile - Future News 24","og_description":"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/","og_site_name":"Future News 24","article_published_time":"2026-07-10T00:00:00+00:00","article_modified_time":"2026-07-12T01:59:31+00:00","og_image":[{"url":"https:\/\/huggingface.co\/blog\/assets\/torch-attention-profile\/thumbnail.png","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/huggingface.co\/blog\/assets\/torch-attention-profile\/thumbnail.png","twitter_misc":{"Written by":"Future News 24","Est. reading time":"21 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Profiling in PyTorch (Half 3): Consideration is all you profile","datePublished":"2026-07-10T00:00:00+00:00","dateModified":"2026-07-12T01:59:31+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/"},"wordCount":4189,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/#primaryimage"},"thumbnailUrl":"https:\/\/huggingface.co\/blog\/assets\/torch-attention-profile\/thumbnail.png","keywords":["Attention","Part","profile","Profiling","PyTorch"],"articleSection":["Developer AI &amp; Open-Source Ecosystem"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/","name":"Profiling in PyTorch (Half 3): Consideration is all you profile - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/#primaryimage"},"thumbnailUrl":"https:\/\/huggingface.co\/blog\/assets\/torch-attention-profile\/thumbnail.png","datePublished":"2026-07-10T00:00:00+00:00","dateModified":"2026-07-12T01:59:31+00:00","description":"We\u2019re on a journey to advance and democratize artificial intelligence through open source and open science.","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/#primaryimage","url":"https:\/\/huggingface.co\/blog\/assets\/torch-attention-profile\/thumbnail.png","contentUrl":"https:\/\/huggingface.co\/blog\/assets\/torch-attention-profile\/thumbnail.png"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/10\/torch-attention-profile\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Profiling in PyTorch (Half 3): Consideration is all you profile"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2213","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=2213"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2213\/revisions"}],"predecessor-version":[{"id":2214,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2213\/revisions\/2214"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/2215"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=2213"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=2213"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=2213"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}