{"id":1874,"date":"2026-06-30T16:00:00","date_gmt":"2026-06-30T16:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/"},"modified":"2026-07-04T17:59:09","modified_gmt":"2026-07-04T17:59:09","slug":"optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/","title":{"rendered":"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">NVIDIA Omniverse NuRec is a neural reconstruction pipeline for constructing high-fidelity 3D representations of real-world environments from multisensor information equivalent to cameras and lidar. It&#8217;s used to reconstruct dynamic scenes captured by autonomous automobile (AV) and robotics platforms into simulation-ready digital environments that may be rendered, replayed, and analyzed inside NVIDIA Omniverse and associated simulation workflows.<\/p>\n<p class=\"wp-block-paragraph\">These reconstructions play a crucial function within the growth of bodily AI and autonomous programs. Engineers can seize a real-world driving or robotics situation, reconstruct the surroundings, after which examine or replay the scene. This permits them to higher perceive mannequin habits, validate notion outcomes, generate artificial viewpoints, or create coaching information for downstream machine studying workflows.<\/p>\n<p class=\"wp-block-paragraph\">NuRec combines neural rendering methods equivalent to Gaussian splatting with GPU-accelerated rendering and simulation pipelines to provide extremely real looking scene reconstructions. Nonetheless, this degree of constancy comes with vital computational price. Reconstruction and rendering workloads contain giant volumes of sensor information, advanced PyTorch-based coaching loops, and extremely specialised CUDA kernels that push GPU sources closely.<\/p>\n<p class=\"wp-block-paragraph\">This submit walks via an instance to showcase the right way to optimize the NuRec neural reconstruction pipeline utilizing NVIDIA Nsight Developer Instruments.<\/p>\n<h2 id=\"solving_performance_optimization_challenges\u00a0\" class=\"wp-block-heading\">Fixing efficiency optimization challenges\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Efficiency is crucial for NuRec workflows as a result of reconstruction turnaround time instantly impacts engineering productiveness. A typical workflow includes figuring out an attention-grabbing or problematic AV run\u2014for instance, a situation the place the notion or planning stack behaved unexpectedly\u2014and launching a reconstruction so engineers can examine the scene as shortly as doable. Ready a number of hours for reconstruction slows iteration and debugging velocity considerably.<\/p>\n<p class=\"wp-block-paragraph\">In the beginning of this optimization effort, reconstructing even comparatively brief captures may take from over an hour to a number of hours relying on the scene and configuration. The group\u2019s long-term purpose is rather more bold: real-time reconstruction efficiency, the place a 30-second seize will be reconstructed in roughly 30 seconds.<\/p>\n<p class=\"wp-block-paragraph\">Efficiency additionally issues past reconstruction itself. As soon as scenes have been reconstructed, rendering-only workflows might generate large numbers of frames for reinforcement studying (RL), artificial information technology (SDG), and large-scale simulation. At this scale, even modest efficiency enhancements can translate instantly into substantial reductions in GPU time and infrastructure price.<\/p>\n<p class=\"wp-block-paragraph\">To sort out these challenges, NVIDIA profiling and optimization instruments had been used, primarily NVIDIA Nsight Methods and NVIDIA Nsight Compute, to investigate the NuRec workload, determine bottlenecks throughout the software program stack, and iteratively optimize each the application-level workflow and the underlying CUDA kernels.<\/p>\n<h2 id=\"profiling_and_optimization_using_nsight_systems\" class=\"wp-block-heading\">Profiling and optimization utilizing Nsight Methods<\/h2>\n<p class=\"wp-block-paragraph\">Nsight Methods is a platform profiling instrument that will help you visualize and perceive the efficiency habits and useful resource utilization of workloads, together with CPU, GPU, storage, networking, and extra. Step one in lots of efficiency optimization workflows is to run an Nsight Methods profile to ascertain a baseline and attempt to determine some preliminary bottlenecks or key areas for enchancment.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">With the purpose of optimizing the coaching loop, we used the Nsight Methods built-in operate help and NVIDIA Instruments Extension SDK (NVTX) included in PyTorch to zoom right into a single iteration of the ahead move proven in Determine 1. The preliminary assumption was that the rendering kernel would take many of the runtime and can be the most effective start line for optimization. Nonetheless, the CUDA HW timeline on the high revealed that almost all of time the GPU was underutilized or not used in any respect. Discover the shortage of blue on the highest row. The applying was additionally utilizing many extra tiny kernels than was anticipated.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec8df60&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec8df60\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"500\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration.webp\" alt=\"Nsight Systems profile timeline screenshot showing nested NVTX ranges for one iteration of the forward pass.&#10;\" class=\"wp-image-119207\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-179x45.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-300x75.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-768x192.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-625x156.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-1536x384.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-645x161.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-500x125.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-160x40.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-362x91.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-440x110.png 440w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-1024x256.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-960x240.png 960w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"500\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration.webp\" alt=\"Nsight Systems profile timeline screenshot showing nested NVTX ranges for one iteration of the forward pass.&#10;\" class=\"lazyload wp-image-119207\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-179x45.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-300x75.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-768x192.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-625x156.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-1536x384.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-645x161.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-500x125.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-160x40.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-362x91.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-440x110.png 440w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-1024x256.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-profile-timeline-one-forward-pass-iteration-960x240.png 960w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 1. Nsight Methods profile timeline for one ahead move iteration<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">After this preliminary realization, it was vital to drill deeper into the phases of the ahead move to determine the place time was being spent and what phases had been underutilizing the GPU. Extra NVTX annotations had been added to the code to delineate numerous phases and features. A brand new profile (Determine 2) confirmed that collect_gaussian_parameters was taking nearly all of the time earlier than rendering even began and is known as a number of instances in every ahead move.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec8ec3a&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec8ec3a\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1524\" height=\"760\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters.webp\" alt=\"Nsight Systems timeline screenshot showing the collect Gaussian parameters function as a large portion of the execution time through NVTX instrumentation.&#10;\" class=\"wp-image-119208\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters.webp 1524w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-179x89.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-300x150.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-768x383.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-625x312.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-645x322.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-500x249.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-160x80.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-362x181.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-221x110.png 221w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-1024x511.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-960x479.png 960w\" sizes=\"(max-width: 1524px) 100vw, 1524px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1524\" height=\"760\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters.webp\" alt=\"Nsight Systems timeline screenshot showing the collect Gaussian parameters function as a large portion of the execution time through NVTX instrumentation.&#10;\" class=\"lazyload wp-image-119208\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters.webp 1524w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-179x89.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-300x150.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-768x383.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-625x312.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-645x322.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-500x249.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-160x80.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-362x181.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-221x110.png 221w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-1024x511.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-collect-gaussian-parameters-960x479.png 960w\" data-sizes=\"(max-width: 1524px) 100vw, 1524px\"\/><figcaption class=\"wp-element-caption\">Determine 2. Figuring out collect_gaussian_parameters as a big portion of execution<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Digging even deeper revealed the interpolate operate taking the plurality of the time (4.148 ms) and calling many small kernels and reminiscence operations that slowed down the GPU, as seen within the backside CUDA API row in Determine 3.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec8fa84&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec8fa84\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"449\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function.webp\" alt=\"Nsight Systems timeline screenshot showing the interpolate function under the collect gaussian parameters function and the many small kernels and memory operations it executes.&#10;\" class=\"wp-image-119212\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-179x40.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-300x67.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-768x173.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-625x140.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-1536x345.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-645x145.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-500x112.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-160x36.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-362x81.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-490x110.png 490w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-1024x230.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-960x216.png 960w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"449\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function.webp\" alt=\"Nsight Systems timeline screenshot showing the interpolate function under the collect gaussian parameters function and the many small kernels and memory operations it executes.&#10;\" class=\"lazyload wp-image-119212\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-179x40.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-300x67.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-768x173.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-625x140.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-1536x345.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-645x145.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-500x112.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-160x36.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-362x81.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-490x110.png 490w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-1024x230.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-interpolate-function-960x216.png 960w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 3. Pinpointing the numerous small kernels of the interpolate operate as an optimization alternative\u00a0<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">We dug into the code below the interpolate operate and targeted on fusing the small kernels and submitting bigger chunks of labor to the GPU. We had been capable of condense all of this work right into a single kernel that diminished the interpolate operate from 4.184 ms to 83.81 us (Determine 4). That is almost a 50x speedup.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec90938&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec90938\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"486\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel.webp\" alt=\"Nsight Systems timeline screenshot showing a single fused kernel under the interpolate function.&#10;\" class=\"wp-image-119210\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-179x44.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-300x73.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-768x187.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-625x152.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-1536x373.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-645x157.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-500x122.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-160x39.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-362x88.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-452x110.png 452w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-1024x249.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-960x233.png 960w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"486\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel.webp\" alt=\"Nsight Systems timeline screenshot showing a single fused kernel under the interpolate function.&#10;\" class=\"lazyload wp-image-119210\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-179x44.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-300x73.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-768x187.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-625x152.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-1536x373.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-645x157.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-500x122.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-160x39.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-362x88.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-452x110.png 452w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-1024x249.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/interpolate-function-single-fused-kernel-960x233.png 960w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 4. Interpolate operate with a single fused kernel on the CUDA API row<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Subsequent we recognized lengthy cudaStreamSynchronize APIs (seen as inexperienced bars on the timeline) that had been delaying the CPU from enqueuing many small kernels whereas the GPU was lively. This resulted in patchy GPU utilization proven within the high CUDA HW row after the synchronize API returned because the small kernels had been scheduled and launched (Determine 5).\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec91604&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec91604\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1608\" height=\"950\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api.webp\" alt=\"Nsight Systems timeline screenshot showing a long CUDA stream synchronize API followed by patchy GPU execution.&#10;\" class=\"wp-image-119214\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api.webp 1608w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-179x106.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-300x177.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-768x454.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-625x369.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-1536x907.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-645x381.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-500x295.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-152x90.png 152w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-362x214.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-186x110.png 186w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-1024x605.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-914x540.png 914w\" sizes=\"(max-width: 1608px) 100vw, 1608px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1608\" height=\"950\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api.webp\" alt=\"Nsight Systems timeline screenshot showing a long CUDA stream synchronize API followed by patchy GPU execution.&#10;\" class=\"lazyload wp-image-119214\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api.webp 1608w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-179x106.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-300x177.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-768x454.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-625x369.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-1536x907.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-645x381.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-500x295.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-152x90.png 152w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-362x214.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-186x110.png 186w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-1024x605.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timline-long-cuda-stream-synch-api-914x540.png 914w\" data-sizes=\"(max-width: 1608px) 100vw, 1608px\"\/><figcaption class=\"wp-element-caption\">Determine 5. Lengthy cudaStreamSynchronize (backside inexperienced row) adopted by patchy GPU execution (high blue row)<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">After eradicating one synchronization level, others down the road would turn out to be the bottleneck. This course of was continued till sufficient had been eliminated that the CPU may effectively enqueue work whereas the GPU was busy. This allowed the tiny kernels to run compactly as a result of they had been not CPU launch-time certain.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec923a9&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec923a9\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"702\" height=\"378\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed.webp\" alt=\"Nsight Systems timeline screenshot showing the previously patchy GPU execution is now more condensed and the cuda stream synchronize API is gone.&#10;\" class=\"wp-image-119215\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed.webp 702w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-179x96.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-300x162.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-625x337.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-645x347.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-500x269.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-160x86.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-362x195.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-204x110.png 204w\" sizes=\"(max-width: 702px) 100vw, 702px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"702\" height=\"378\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed.webp\" alt=\"Nsight Systems timeline screenshot showing the previously patchy GPU execution is now more condensed and the cuda stream synchronize API is gone.&#10;\" class=\"lazyload wp-image-119215\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed.webp 702w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-179x96.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-300x162.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-625x337.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-645x347.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-500x269.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-160x86.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-362x195.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-timeline-after-syncronization-points-removed-204x110.png 204w\" data-sizes=\"(max-width: 702px) 100vw, 702px\"\/><figcaption class=\"wp-element-caption\">Determine 6. Compact GPU utilization (high blue row) after synchronization factors eliminated<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Lowering the time spent accumulating the parameters and eradicating synchronization factors that had been inflicting bottlenecks enabled digging into some kernel optimizations. Nsight Methods allows you to determine which kernels are the most well liked. The renderBackward kernel was clearly the highest candidate on this case.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec93040&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec93040\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"699\" height=\"144\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time.webp\" alt=\"Nsight Systems screenshot showing the top kernels in the CUDA hardware row of the timeline.\" class=\"wp-image-119216\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time.webp 699w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-179x37.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-300x62.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-625x129.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-645x133.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-500x103.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-160x33.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-362x75.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-534x110.png 534w\" sizes=\"(max-width: 699px) 100vw, 699px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"699\" height=\"144\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time.webp\" alt=\"Nsight Systems screenshot showing the top kernels in the CUDA hardware row of the timeline.\" class=\"lazyload wp-image-119216\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time.webp 699w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-179x37.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-300x62.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-625x129.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-645x133.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-500x103.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-160x33.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-362x75.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-systems-top-kernels-by-execution-time-534x110.png 534w\" data-sizes=\"(max-width: 699px) 100vw, 699px\"\/><figcaption class=\"wp-element-caption\">Determine 7. Prime kernels by execution time in Nsight Methods<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"kernel_optimization_using_nsight_compute\" class=\"wp-block-heading\">Kernel optimization utilizing Nsight Compute<\/h2>\n<p class=\"wp-block-paragraph\">Nsight Compute is the most effective instrument for profiling and optimizing particular person kernels. It will probably robotically replay kernels to gather giant quantities of efficiency information at very fantastic granularities utilizing numerous sorts of {hardware} counters, software program patching, and instrumentation. It features a built-in rule system and guided evaluation to assist customers determine and perceive points.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The renderBackward kernel is utilized in each digicam and lidar information processing. Profiling a number of cases of this kernel with Nsight Compute revealed that it has solely ~15% occupancy and the habits and useful resource necessities of this kernel differ considerably relying on which of those inputs is being processed.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The longest three renderBackward kernels are from lidar information and the opposite three are from digicam information. Regardless of these variations, each had been allocating 167 registers per thread (Determine 8).<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec93f6f&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec93f6f\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1600\" height=\"239\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage.webp\" alt=\"Screenshot of Nsight Compute summary page showing the top six kernels, durations, and resource allocations.&#10;\" class=\"wp-image-119217\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage.webp 1600w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-179x27.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-300x45.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-768x115.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-625x93.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-1536x229.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-645x96.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-500x75.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-160x24.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-362x54.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-736x110.png 736w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-1024x153.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-960x143.png 960w\" sizes=\"(max-width: 1600px) 100vw, 1600px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1600\" height=\"239\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage.webp\" alt=\"Screenshot of Nsight Compute summary page showing the top six kernels, durations, and resource allocations.&#10;\" class=\"lazyload wp-image-119217\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage.webp 1600w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-179x27.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-300x45.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-768x115.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-625x93.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-1536x229.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-645x96.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-500x75.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-160x24.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-362x54.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-736x110.png 736w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-1024x153.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-kernel-performance-resource-usage-960x143.png 960w\" data-sizes=\"(max-width: 1600px) 100vw, 1600px\"\/><figcaption class=\"wp-element-caption\">Determine 8. Six profiled cases of the renderBackward kernel in Nsight Compute<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Setting the highest lidar kernel as an Nsight Compute baseline and evaluating a digicam kernel robotically revealed that whereas each had the overwhelming majority of accesses in shared reminiscence, the digicam kernels had been making ~75% fewer requests despite the fact that each lidar and digicam cases of the kernel had been allocating the identical quantity of shared reminiscence per block statically.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec94e05&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec94e05\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1288\" height=\"291\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats.webp\" alt=\"Nsight Compute memory statistics table showing the difference in shared memory accesses between lidar and camera kernels.&#10;\" class=\"wp-image-119218\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats.webp 1288w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-179x40.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-300x68.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-768x174.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-625x141.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-645x146.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-500x113.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-160x36.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-362x82.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-487x110.png 487w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-1024x231.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-960x217.png 960w\" sizes=\"(max-width: 1288px) 100vw, 1288px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1288\" height=\"291\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats.webp\" alt=\"Nsight Compute memory statistics table showing the difference in shared memory accesses between lidar and camera kernels.&#10;\" class=\"lazyload wp-image-119218\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats.webp 1288w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-179x40.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-300x68.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-768x174.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-625x141.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-645x146.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-500x113.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-160x36.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-362x82.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-487x110.png 487w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-1024x231.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-memory-stats-960x217.png 960w\" data-sizes=\"(max-width: 1288px) 100vw, 1288px\"\/><figcaption class=\"wp-element-caption\">Determine 9. Distinction in shared reminiscence accesses between lidar and digicam information kernels<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Noting these habits variations between whether or not the renderBackward kernel was used for digicam or lidar information, and the truth that register and shared reminiscence allocations had been static and equivalent for each, the following step was to attempt splitting the kernel relying on whether or not it was processing digicam or lidar information.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">For every model of the kernel, the group experimented and tuned register allocations with the launch_bounds qualifier and the quantity of shared reminiscence we had been allocating per block. The cudaFuncSetCacheConfig runtime API was used to set the choice of each kernels to have a bigger shared reminiscence and smaller L1 cache.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">After this testing and optimization, the lidar and digicam kernels decreased their register allocation wants from 167 to 64 and 128 respectively and each had been capable of run effectively with about half of the initially allotted shared reminiscence. This improved occupancy from ~15% to between 30-50% and general runtime considerably, with the longest lidar kernel reducing from 31 ms to 18 ms.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec95d58&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec95d58\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1866\" height=\"242\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel.webp\" alt=\"Nsight Compute summary page showing the six kernels\u2019 performance and resource usage after separating lidar from camera processing.&#10;\" class=\"wp-image-119219\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel.webp 1866w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-179x23.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-300x39.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-768x100.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-625x81.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-1536x199.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-645x84.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-500x65.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-160x21.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-362x47.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-848x110.png 848w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-1024x133.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-960x125.png 960w\" sizes=\"(max-width: 1866px) 100vw, 1866px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1866\" height=\"242\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel.webp\" alt=\"Nsight Compute summary page showing the six kernels\u2019 performance and resource usage after separating lidar from camera processing.&#10;\" class=\"lazyload wp-image-119219\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel.webp 1866w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-179x23.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-300x39.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-768x100.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-625x81.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-1536x199.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-645x84.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-500x65.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-160x21.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-362x47.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-848x110.png 848w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-1024x133.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-summary-renderbackward-kernel-960x125.png 960w\" data-sizes=\"(max-width: 1866px) 100vw, 1866px\"\/><figcaption class=\"wp-element-caption\">Determine 10. Kernel efficiency and configurations after splitting lidar and digicam processing<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">However there may be nonetheless room for enchancment. The following situation recognized, which is being labored on on the time of publication, is long-tail results within the kernel brought on by a workload imbalance. This may be seen within the PM Sampling part of Nsight Compute (Determine 11). The primary half of the kernel reveals a median of 32 lively warps that start to taper off and for the final a number of milliseconds there may be lower than one lively warp per cycle. Ideally, all of the warps can be lively for the whole thing of the kernel.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a4949ec96d4c&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a4949ec96d4c\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"254\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance.webp\" alt=\"Nsight Compute PM Sampling section showing a long tail of active warps indicating a load imbalance issue.&#10;\" class=\"wp-image-119220\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-179x23.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-300x38.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-768x98.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-625x79.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-1536x195.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-645x82.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-500x64.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-160x20.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-362x46.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-866x110.png 866w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-1024x130.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-960x122.png 960w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"254\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance.webp\" alt=\"Nsight Compute PM Sampling section showing a long tail of active warps indicating a load imbalance issue.&#10;\" class=\"lazyload wp-image-119220\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-179x23.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-300x38.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-768x98.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-625x79.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-1536x195.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-645x82.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-500x64.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-160x20.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-362x46.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-866x110.png 866w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-1024x130.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/nsight-compute-pm-sampling-load-imbalance-960x122.png 960w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 11. Lengthy-tail impact proven in Nsight Compute indicating a load imbalance<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Efficiency evaluation and optimization is an iterative course of that consists of operating a profile, figuring out an issue, fixing it, and beginning once more. You should utilize instruments like Nsight Methods and Nsight Compute to make this whole course of simpler for creating and optimizing on NVIDIA GPUs. Each instruments are free\u2014obtain Nsight Methods and Nsight Compute and check out them with your personal use case. In case you have questions or need to share what you discover, depart a touch upon the NVIDIA Developer Boards.\u00a0<\/p>\n<h3 id=\"acknowledgments\" class=\"wp-block-heading\">Acknowledgments<\/h3>\n<p class=\"wp-block-paragraph\">Particular because of NVIDIA contributors Francois Trudel, Joey Lai, and Rodolfo Lima.\u00a0<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>NVIDIA Omniverse NuRec is a neural reconstruction pipeline for constructing high-fidelity 3D representations of real-world environments from multisensor information equivalent to cameras and lidar. It&#8217;s used to reconstruct dynamic scenes captured by autonomous automobile (AV) and robotics platforms into simulation-ready digital environments that may be rendered, replayed, and analyzed inside NVIDIA Omniverse and associated simulation [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1876,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/av.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[229,2359,2361,81,2053,1139,2360,56],"class_list":["post-1874","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-developer","tag-neural","tag-nsight","tag-nvidia","tag-optimizing","tag-pipeline","tag-reconstruction","tag-tools"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments - Future News 24<\/title>\n<meta name=\"description\" content=\"NVIDIA Omniverse NuRec is a neural reconstruction pipeline for building high&#x2d;fidelity 3D representations of real&#x2d;world environments from multisensor data such&#8230;\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments - Future News 24\" \/>\n<meta property=\"og:description\" content=\"NVIDIA Omniverse NuRec is a neural reconstruction pipeline for building high&#x2d;fidelity 3D representations of real&#x2d;world environments from multisensor data such&#8230;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-30T16:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-04T17:59:09+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/av.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/av.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"8 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments\",\"datePublished\":\"2026-06-30T16:00:00+00:00\",\"dateModified\":\"2026-07-04T17:59:09+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/\"},\"wordCount\":1633,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/av.webp\",\"keywords\":[\"Developer\",\"Neural\",\"Nsight\",\"NVIDIA\",\"Optimizing\",\"pipeline\",\"Reconstruction\",\"Tools\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/\",\"name\":\"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/av.webp\",\"datePublished\":\"2026-06-30T16:00:00+00:00\",\"dateModified\":\"2026-07-04T17:59:09+00:00\",\"description\":\"NVIDIA Omniverse NuRec is a neural reconstruction pipeline for building high&#x2d;fidelity 3D representations of real&#x2d;world environments from multisensor data such&#8230;\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/av.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/av.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments - Future News 24","description":"NVIDIA Omniverse NuRec is a neural reconstruction pipeline for building high&#x2d;fidelity 3D representations of real&#x2d;world environments from multisensor data such&#8230;","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/","og_locale":"en_US","og_type":"article","og_title":"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments - Future News 24","og_description":"NVIDIA Omniverse NuRec is a neural reconstruction pipeline for building high&#x2d;fidelity 3D representations of real&#x2d;world environments from multisensor data such&#8230;","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/","og_site_name":"Future News 24","article_published_time":"2026-06-30T16:00:00+00:00","article_modified_time":"2026-07-04T17:59:09+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/av.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/av.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"8 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments","datePublished":"2026-06-30T16:00:00+00:00","dateModified":"2026-07-04T17:59:09+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/"},"wordCount":1633,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/av.webp","keywords":["Developer","Neural","Nsight","NVIDIA","Optimizing","pipeline","Reconstruction","Tools"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/","name":"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/av.webp","datePublished":"2026-06-30T16:00:00+00:00","dateModified":"2026-07-04T17:59:09+00:00","description":"NVIDIA Omniverse NuRec is a neural reconstruction pipeline for building high&#x2d;fidelity 3D representations of real&#x2d;world environments from multisensor data such&#8230;","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/av.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/av.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/optimizing-a-neural-reconstruction-pipeline-using-nvidia-nsight-developer-tools\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Optimizing a Neural Reconstruction Pipeline Utilizing NVIDIA Nsight Developer Instruments"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1874","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=1874"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1874\/revisions"}],"predecessor-version":[{"id":1875,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1874\/revisions\/1875"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/1876"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=1874"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=1874"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=1874"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}