{"id":1464,"date":"2026-06-24T16:30:00","date_gmt":"2026-06-24T16:30:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/"},"modified":"2026-06-25T09:59:25","modified_gmt":"2026-06-25T09:59:25","slug":"accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/","title":{"rendered":"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">An more and more widespread design sample for autonomous autos (AVs), robotics, and spatial AI techniques is fowl\u2019s-eye-view (BEV) notion. BEV fashions venture multicamera picture options right into a shared top-down grid, offering downstream notion and planning modules with a standard spatial format for reasoning about lanes, autos, pedestrians, and free house.<\/p>\n<p class=\"wp-block-paragraph\">A key operation on this pipeline is BEV pooling, which gathers picture options, weights them with depth data, and scatter-reduces them into BEV grid cells. For builders, the sensible worth of BEV notion is that it converts many camera-specific views into one spatially constant illustration of the scene. As an alternative of reasoning individually over every digicam picture, downstream modules can function on a unified top-down function map aligned to the world across the automobile or robotic. BEV pooling is the step that makes this illustration usable in actual time: it turns depth-aware picture options right into a compact BEV tensor that may feed detection, occupancy, trajectory prediction, mapping, and planning workloads.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Conceptually, that is easy. In deployment, nonetheless, BEV pooling can turn into a latency bottleneck as a result of it combines irregular reminiscence entry, repeated index reads, scatter-reduce conduct, and GPU-specific cache results.<\/p>\n<p class=\"wp-block-paragraph\">This put up makes use of BEVPoolV3 as a case examine in optimizing BEV pooling and different gather- or scatter-heavy operators for NVIDIA GPUs. It walks by way of a sensible workflow you possibly can apply to your workloads: classify the reminiscence regime, take away redundant scatter visitors, map the kernel implementation to the goal GPU, and validate the lively bottleneck with NVIDIA Nsight Compute. The efficiency outcomes present why this workflow issues: the identical BEV pooling operator can require completely different optimization methods relying on whether or not the working set is DRAM-bound or largely L2-resident.<\/p>\n<h2 id=\"how_does_bevpoolv3_reduce_bev_pooling_latency_on_nvidia_rtx_gpus\u00a0\" class=\"wp-block-heading\">How does BEVPoolV3 cut back BEV pooling latency on NVIDIA RTX GPUs?\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">Prior work has already made necessary progress. BEVPoolV2, known as V2 on this put up, launched an environment friendly deployment-oriented BEV pooling formulation for BEVDet-style fashions. CUDA-BEVFusion contains bevpool_half_pack10_kernel, referred to right here as V2+DO, which makes use of depth-outer traversal to take away a lot of the V2 repeated tile-outer index loading.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">BEVPoolV3 continues this optimization route with 4 further adjustments: lowered duplicate depth masses, a five-array INT32 scatter map, precomputed indices that take away runtime integer division, and interval-owned output writes.<\/p>\n<p class=\"wp-block-paragraph\">This put up makes use of BEVPoolV3 as a case examine in the right way to optimize BEV pooling and different gather- or scatter-heavy operators for NVIDIA GPUs. You&#8217;ll learn to classify a BEV pooling workload by reminiscence regime, determine redundant scatter visitors, map the kernel implementation to the goal GPU, and validate the lively bottleneck with Nsight Compute. The efficiency outcomes on two NVIDIA RTX GPUs present why this workflow issues: the identical BEV pooling algorithm will be DRAM-bound on one GPU and largely L2-resident on one other, requiring completely different optimization selections.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The analysis compares two NVIDIA RTX GPUs that signify completely different reminiscence regimes: NVIDIA RTX A6000, an NVIDIA Ampere SM86 GPU with a 6 MB L2 cache and no native FP8 ISA, and NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Version, an NVIDIA Blackwell SM120 GPU with a 128 MB L2 cache and native FP8 help. The canonical config used right here is derived from actual nuScenes samples and accommodates about 209K scatter factors, 80 function channels, and a 49 MB BEV pooling working set. That working set exceeds RTX A6000 L2 cache however matches inside RTX PRO 6000 Blackwell Max-Q L2 cache, making RTX A6000 DRAM-bound and RTX PRO 6000 Blackwell Max-Q largely L2-resident after the preliminary fill.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a3cfbfc7e2a7&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a3cfbfc7e2a7\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"979\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1.webp\" alt=\"Four-panel diagram titled &quot;Converting multi-camera image features into a top-down representation for downstream perception and planning.&quot; Panel 1 shows six camera views (front, left, right, left rear, right rear, rear) around a vehicle, producing multi-camera image features. Panel 2 lifts those features into a 3D grid of depth-weighted points. Panel 3 scatters those points into a top-down BEV grid (the BEV pooling step). Panel 4 shows the resulting BEV feature map feeding three downstream tasks: detection, occupancy, and planning. A green callout reads &quot;Why it matters: BEV pooling is a key deployment bottleneck in camera-based AV perception.&quot;\" class=\"wp-image-118950\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-179x88.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-300x147.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-768x376.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-625x306.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-1536x752.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-645x316.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-500x245.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-160x78.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-362x177.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-225x110.png 225w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-1024x501.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-960x470.png 960w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"979\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1.webp\" alt=\"Four-panel diagram titled &quot;Converting multi-camera image features into a top-down representation for downstream perception and planning.&quot; Panel 1 shows six camera views (front, left, right, left rear, right rear, rear) around a vehicle, producing multi-camera image features. Panel 2 lifts those features into a 3D grid of depth-weighted points. Panel 3 scatters those points into a top-down BEV grid (the BEV pooling step). Panel 4 shows the resulting BEV feature map feeding three downstream tasks: detection, occupancy, and planning. A green callout reads &quot;Why it matters: BEV pooling is a key deployment bottleneck in camera-based AV perception.&quot;\" class=\"lazyload wp-image-118950\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-179x88.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-300x147.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-768x376.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-625x306.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-1536x752.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-645x316.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-500x245.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-160x78.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-362x177.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-225x110.png 225w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-1024x501.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/bev-pooling-convert-multicamera-image-features-1-960x470.png 960w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 1. BEV pooling lifts multicamera picture options with depth data and scatter-reduces them right into a shared top-down illustration for detection, occupancy prediction, and planning<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Within the canonical config, the V2-style NVIDIA TensorRT plugin path takes 274.0 \u00b5s on RTX PRO 6000 Blackwell Max-Q. BEVPoolV3 reduces that to 17.3 \u00b5s in FP16 and 16.4 \u00b5s in FP8. On RTX A6000, the DRAM-adapted BEVPoolV3 FP16 path reaches 90.0 \u00b5s. Past the speedup, this put up exhibits a repeatable workflow for optimizing scatter-reduce kernels: classify the working set, take away redundant reminiscence visitors, match the launch form to the goal GPU, and validate the outcome with Nsight Compute.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a3cfbfc7f0e9&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a3cfbfc7f0e9\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"875\" height=\"527\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1.webp\" alt=\"Bar chart of speedup over V2 FP16 at the canonical config (TRT 100-iteration median). Bars left to right: V2 FP16 baseline at 1.00\u00d7 (274.0 \u00b5s), RTX A6000 V3 FP16 at 19.31\u00d7 (90.0 \u00b5s), RTX PRO 6000 Blackwell Max-QV2+DO FP16 at 7.25\u00d7 (37.8 \u00b5s, shown in grey), RTX PRO 6000 Blackwell Max-Q V3 FP16 at 15.84\u00d7  (17.3 \u00b5s), and RTX PRO 6000 Blackwell Max-Q V3 FP8 at 16.71\u00d7 (16.4 \u00b5s). The three V3 bars are green; the V2+DO baseline is grey. A dashed horizontal line marks the 1\u00d7 baseline.&#10;\" class=\"wp-image-119004\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1.webp 875w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-179x108.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-300x181.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-768x463.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-625x376.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-645x388.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-498x300.png 498w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-149x90.png 149w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-362x218.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-183x110.png 183w\" sizes=\"(max-width: 875px) 100vw, 875px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"875\" height=\"527\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1.webp\" alt=\"Bar chart of speedup over V2 FP16 at the canonical config (TRT 100-iteration median). Bars left to right: V2 FP16 baseline at 1.00\u00d7 (274.0 \u00b5s), RTX A6000 V3 FP16 at 19.31\u00d7 (90.0 \u00b5s), RTX PRO 6000 Blackwell Max-QV2+DO FP16 at 7.25\u00d7 (37.8 \u00b5s, shown in grey), RTX PRO 6000 Blackwell Max-Q V3 FP16 at 15.84\u00d7  (17.3 \u00b5s), and RTX PRO 6000 Blackwell Max-Q V3 FP8 at 16.71\u00d7 (16.4 \u00b5s). The three V3 bars are green; the V2+DO baseline is grey. A dashed horizontal line marks the 1\u00d7 baseline.&#10;\" class=\"lazyload wp-image-119004\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1.webp 875w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-179x108.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-300x181.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-768x463.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-625x376.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-645x388.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-498x300.png 498w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-149x90.png 149w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-362x218.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-speedup-bevpoolv3-1-183x110.png 183w\" data-sizes=\"(max-width: 875px) 100vw, 875px\"\/><figcaption class=\"wp-element-caption\">Determine 2. Canonical TensorRT plugin path speedup over V2 FP16. On RTX A6000, V3 FP16 reaches 19.31x over V2. On RTX PRO 6000 Blackwell Max-Q, V3 FP16 reaches 15.84x over V2 and V3 FP8 reaches 16.71x over V2<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"prerequisites\u00a0\" class=\"wp-block-heading\">Stipulations\u00a0<\/h2>\n<p class=\"wp-block-paragraph\">This put up discusses CUDA kernel conduct, TensorRT plugin integration, and GPU profiling within the context of BEV pooling. Useful conditions embody:<\/p>\n<p>CUDA kernel ideas akin to warp scheduling, atomics, vectorized world masses, and DRAM\/L2\/L1 cache conduct<\/p>\n<p>TensorRT plugin integration, particularly the IPluginV3 interface<\/p>\n<p>Nsight Compute profiling for validating reminiscence conduct, occupancy, and instruction-issue bottlenecks<\/p>\n<p>The BEV-pooling kernel in CUDA-BEVFusion because the prior depth-outer reference implementation<\/p>\n<p class=\"wp-block-paragraph\">For associated background data, see the CUDA C++ Programming Information, TensorRT plugin documentation, TensorRT samples, and Nsight Compute Profiling Information.<\/p>\n<h2 id=\"classify_the_memory_regime\" class=\"wp-block-heading\">Classify the reminiscence regime<\/h2>\n<p class=\"wp-block-paragraph\">Step one is to categorise whether or not the BEV-pooling working set matches in L2. Within the canonical config, the principle arrays complete about 49 MB, dominated by function information and output. That single quantity determines the reminiscence regime: it&#8217;s bigger than the RTX A6000 6 MB L2 cache, however smaller than RTX PRO 6000 Blackwell Max-Q 128 MB L2 cache.<\/p>\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a3cfbfc80340&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a3cfbfc80340\" class=\"wp-block-image size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"952\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1.webp\" alt=\"Two stacked horizontal bar charts comparing the ~49 MB canonical BEV-pooling working set against each GPU's L2 cache capacity. Top: RTX A6000 with 6 MB L2 \u2014 the working set (green bar) extends far past the cache capacity line, labeled &quot;Does not fit \u2192 DRAM-bound path.&quot; Bottom: RTX PRO 6000 Blackwell Max-Q with 128 MB L2 \u2014 the working set fills only a small fraction of the bar, labeled &quot;Fits \u2192 L2-resident path.&quot; Subtitle: &quot;Derived from real nuScenes samples, C=80, ~209K scatter points.&quot;&#10;\" class=\"wp-image-118952\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-179x85.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-300x143.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-768x366.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-625x298.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-1536x732.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-645x307.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-500x238.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-160x76.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-362x172.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-231x110.png 231w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-1024x488.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-960x457.png 960w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"952\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1.webp\" alt=\"Two stacked horizontal bar charts comparing the ~49 MB canonical BEV-pooling working set against each GPU's L2 cache capacity. Top: RTX A6000 with 6 MB L2 \u2014 the working set (green bar) extends far past the cache capacity line, labeled &quot;Does not fit \u2192 DRAM-bound path.&quot; Bottom: RTX PRO 6000 Blackwell Max-Q with 128 MB L2 \u2014 the working set fills only a small fraction of the bar, labeled &quot;Fits \u2192 L2-resident path.&quot; Subtitle: &quot;Derived from real nuScenes samples, C=80, ~209K scatter points.&quot;&#10;\" class=\"lazyload wp-image-118952\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-179x85.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-300x143.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-768x366.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-625x298.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-1536x732.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-645x307.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-500x238.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-160x76.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-362x172.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-231x110.png 231w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-1024x488.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/canonical-bev-pooling-working-set-1-960x457.png 960w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 3. Classifying the canonical BEV-pooling working set by L2 capability. The canonical config derived from actual nuScenes samples has a working set of about 49 MB. This exceeds the 6 MB L2 cache on RTX A6000, so the kernel follows a DRAM-bound path. The identical working set matches contained in the 128 MB L2 cache on RTX PRO 6000 Blackwell Max-Q, so the kernel is basically L2-resident after the preliminary fill. Observe that the diagram is conceptual and never drawn to precise scale<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">This match\/no-fit resolution adjustments the optimization goal. On RTX A6000, function gathers and output visitors spill past L2, so the small-L2 path prioritizes byte discount and cache-streaming output shops. On RTX PRO 6000 Blackwell Max-Q, the canonical working set matches in L2, so the large-L2 path shifts towards instruction effectivity, occupancy, precomputed indices, vectorized masses, and FP8 specialization.<\/p>\n<h2 id=\"remove_redundant_scatter_traffic\" class=\"wp-block-heading\">Take away redundant scatter visitors<\/h2>\n<p class=\"wp-block-paragraph\">The BEV scatter-reduce will be summarized as:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nout[ranks_bev[t], c] += depth[ranks_depth[t]] * feat[ranks_feat[t], c];\n<\/div>\n<p class=\"wp-block-paragraph\">BEVPoolV2 iterates over channel tiles outdoors the scatter loop. For C=80 and an 8-channel tile, the identical scatter indices are loaded 10 instances. That produces roughly 25.1 MB of index visitors for indices that solely want 2.51 MB when learn as soon as. A depth-outer loop order fixes most of that drawback by iterating over every BEV interval first and accumulating all channels for that interval in a single move.<\/p>\n<p class=\"wp-block-paragraph\">BEVPoolV3 extends the depth-outer optimization route utilized in CUDA-BEVFusion bevpool_half_pack10_kernel, referred to right here as V2+DO. V2+DO is a helpful baseline as a result of it already removes the repeated tile-outer index masses in BEVPoolV2 and demonstrates the worth of interval-based traversal. BEVPoolV3 retains that route and provides 4 implementation adjustments that enhance portability and efficiency throughout GPU reminiscence regimes: lowered duplicate depth masses inside every interval; a five-array INT32 scatter map \u00b5sing ranks_depth, ranks_feat, ranks_bev, interval_starts, and interval_lengths; precomputed specific indices that take away runtime integer division; and interval-owned output writes that keep away from atomics relative to the V2-style path.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a3cfbfc813a9&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a3cfbfc813a9\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"1125\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart.webp\" alt=\"Stacked bar chart of analytical memory traffic in MB at the canonical config (N=209K,  C=80). Four bars left to right: V2 FP16 \u224877 MB (largest scatter-indices segment, plus depth, feat, and intervals+output), V2+DO FP16 \u224851 MB (scatter-indices collapse, feat dominates), V3 FP16 \u224850 MB (slightly smaller scatter-indices, feat still dominant), V3 FP8 \u224826 MB (feat and intervals+output both halve). Stack segments from bottom to top: scatter indices (darkest green), depth, feat, intervals plus output (lightest green).&#10;\" class=\"wp-image-119001\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-179x101.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-300x169.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-768x432.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-625x352.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-1536x864.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-645x363.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-660x370.png 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-500x281.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-160x90.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-362x204.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-195x110.png 195w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-1024x576.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-960x540.png 960w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"1125\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart.webp\" alt=\"Stacked bar chart of analytical memory traffic in MB at the canonical config (N=209K,  C=80). Four bars left to right: V2 FP16 \u224877 MB (largest scatter-indices segment, plus depth, feat, and intervals+output), V2+DO FP16 \u224851 MB (scatter-indices collapse, feat dominates), V3 FP16 \u224850 MB (slightly smaller scatter-indices, feat still dominant), V3 FP8 \u224826 MB (feat and intervals+output both halve). Stack segments from bottom to top: scatter indices (darkest green), depth, feat, intervals plus output (lightest green).&#10;\" class=\"lazyload wp-image-119001\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-179x101.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-300x169.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-768x432.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-625x352.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-1536x864.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-645x363.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-660x370.png 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-500x281.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-160x90.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-362x204.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-195x110.png 195w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-1024x576.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/analytical-memory-traffic-chart-960x540.png 960w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 4. V2+DO removes most redundant index visitors. V3 FP16 additional reduces aligned scatter-map overhead and instruction strain. V3 FP8 halves function and output bytes, which helps most when the working set is L2-resident<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">The five-array scatter map is very necessary on large-L2 GPUs. Packing (ranks_depth, ranks_feat, ranks_bev) into an int3 array offers a 12-byte report. That format is inconvenient for aligned reminiscence transactions and doesn&#8217;t map cleanly to a 16-byte LDG.128 load. Separate INT32 arrays let adjoining threads merge aligned masses and keep away from area coupling. The whole logical bytes might look comparable, however the instruction stream is way cleaner.<\/p>\n<h2 id=\"implement_interval-owned_scatter-reduce\" class=\"wp-block-heading\">Implement interval-owned scatter-reduce<\/h2>\n<p class=\"wp-block-paragraph\">In manufacturing, BEVPoolV3 makes use of a number of specialised kernels, however the core implementation concept is less complicated to grasp as a small logic sketch. The scatter map is ready forward of time, every BEV interval is assigned to 1 proprietor, the proprietor walks the factors in that interval, accumulates the related function channels, and writes the output as soon as.<\/p>\n<p class=\"wp-block-paragraph\">This construction removes the inner-loop decoding work that seems when the scatter map is packed right into a single report. As an alternative of reconstructing indices at runtime, the kernel reads specific arrays akin to ranks_depth, ranks_feat, ranks_bev, interval_starts, and interval_lengths.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\n\/\/ 1. Use 5 precomputed scatter arrays.<br \/>\n\/\/ 2. Learn specific indices instantly, with no runtime index division.<br \/>\n\/\/ 3. Let one interval proprietor accumulate the output cell.<br \/>\n\/\/ 4. Load every depth worth as soon as per scatter level within the proprietor loop.<\/p>\n<p>for every interval iv in parallel:<br \/>\n    begin  = interval_starts[iv]<br \/>\n    size = interval_lengths[iv]<br \/>\n    bev    = ranks_bev[start]<\/p>\n<p>    acc[channel_tile] = 0<\/p>\n<p>    for offset in 0 .. size &#8211; 1:<br \/>\n        t        = begin + offset<br \/>\n        d        = depth[ranks_depth[t]]<br \/>\n        feat_row = ranks_feat[t]<\/p>\n<p>        for c in channel_tile:<br \/>\n            acc[c] += d * feat[feat_row, c]<\/p>\n<p>    out[bev, channel_tile] = acc\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">This code sketch captures the widespread BEVPoolV3 construction: the scatter map is specific, runtime index decoding is eliminated, depth is loaded within the interval proprietor loop, and every output cell is written as soon as after native accumulation.<\/p>\n<p class=\"wp-block-paragraph\">The manufacturing kernels specialize this construction for the goal reminiscence regime. On small-L2 GPUs akin to RTX A6000, the implementation prioritizes byte discount, FP16 half2 accumulation, and cache-streaming output shops so the output tensor doesn&#8217;t evict helpful index information from L2. On large-L2 GPUs akin to RTX PRO 6000 Blackwell Max-Q, the implementation first matches a high-occupancy launch envelope, then reduces instruction overhead with precomputed indices, vectorized index masses, and FP8-specialized interior loops the place the working set is L2-resident.<\/p>\n<p class=\"wp-block-paragraph\">The algorithmic invariant stays the identical: personal the interval, keep away from runtime index decoding, accumulate regionally, and write as soon as. The architecture-specific work adjustments how that invariant is carried out, not what the BEV-pooling operator computes.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a3cfbfc82641&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a3cfbfc82641\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"824\" height=\"484\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt.webp\" alt=\"Grouped bar chart of TRT-11 latency on RTX PRO 6000 Blackwell Max-Q in microseconds (log scale, 100-iteration median) across six configurations: small-100K, canonical-209K, large-500K, xlarge-1M, c128-209K, c256-209K. Each config has four bars: V2 FP16 (light grey, tallest, ranging from ~140 \u00b5s to ~1700 \u00b5s), V2+DO FP16 (dark grey), V3 FP16 (light green), and V3 FP8 (NVIDIA green, shortest in every config). V3 FP8 is fastest on every shape; the V2 vs V3 gap widens at larger N and wider channels.&#10;\" class=\"wp-image-119006\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt.webp 824w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-179x105.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-300x176.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-768x451.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-625x367.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-645x379.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-500x294.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-153x90.png 153w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-362x213.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-187x110.png 187w\" sizes=\"(max-width: 824px) 100vw, 824px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"824\" height=\"484\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt.webp\" alt=\"Grouped bar chart of TRT-11 latency on RTX PRO 6000 Blackwell Max-Q in microseconds (log scale, 100-iteration median) across six configurations: small-100K, canonical-209K, large-500K, xlarge-1M, c128-209K, c256-209K. Each config has four bars: V2 FP16 (light grey, tallest, ranging from ~140 \u00b5s to ~1700 \u00b5s), V2+DO FP16 (dark grey), V3 FP16 (light green), and V3 FP8 (NVIDIA green, shortest in every config). V3 FP8 is fastest on every shape; the V2 vs V3 gap widens at larger N and wider channels.&#10;\" class=\"lazyload wp-image-119006\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt.webp 824w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-179x105.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-300x176.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-768x451.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-625x367.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-645x379.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-500x294.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-153x90.png 153w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-362x213.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/latency-analysis-tensorrt-187x110.png 187w\" data-sizes=\"(max-width: 824px) 100vw, 824px\"\/><figcaption class=\"wp-element-caption\">Determine 5. RTX PRO 6000 Blackwell Max-Q TensorRT latency throughout six BEV pooling configurations. V3 FP8 is quickest on each configuration, with bigger positive aspects on wider channel counts<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Absolutely the latency outcomes on RTX PRO 6000 Blackwell Max-Q present how the large-L2 path behaves throughout completely different level counts and channel widths. The identical optimization sample additionally holds on the RTX A6000 DRAM-bound path when measured as speedup over the V2 FP16 baseline. On RTX A6000, the DRAM-adapted V3 FP16 path reaches speedups of 11s to 22x over V2 throughout the examined configurations. On RTX PRO 6000 Blackwell Max-Q, V3 FP8 reaches speedups of 11x to 42x over V2, with the biggest positive aspects showing at bigger level counts and wider channel configurations.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a3cfbfc834bc&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a3cfbfc834bc\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"832\" height=\"457\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1.webp\" alt=\"Grouped bar chart of speedup over V2 FP16 across six configs (small, canonical, large, xlarge, wide-c128, wide-c256). Each group has four bars: RTX PRO 6000 Blackwell Max-Q V2+DO FP16\/V2 (grey), RTX PRO 6000 Blackwell Max-Q V3 FP16\/V2 (dark green), RTX PRO 6000 Blackwell Max-Q V3 FP8\/V2 (light green), RTX A6000 V3 FP16\/V2 (medium green). RTX PRO 6000 Blackwell Max-Q V3 FP8 reaches the highest speedup on every RTX PRO 6000 Blackwell Max-Q config, peaking at 42\u00d7 on xlarge and 40\u00d7 on wide-c256. RTX A6000 V3 FP16 peaks at the small\/canonical end (~21\u00d7\/19\u00d7) and declines on larger configs where it becomes DRAM-bound.\" class=\"wp-image-119008\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1.webp 832w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-179x98.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-300x165.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-768x422.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-625x343.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-645x354.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-500x275.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-160x88.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-362x199.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-200x110.png 200w\" sizes=\"(max-width: 832px) 100vw, 832px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"832\" height=\"457\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1.webp\" alt=\"Grouped bar chart of speedup over V2 FP16 across six configs (small, canonical, large, xlarge, wide-c128, wide-c256). Each group has four bars: RTX PRO 6000 Blackwell Max-Q V2+DO FP16\/V2 (grey), RTX PRO 6000 Blackwell Max-Q V3 FP16\/V2 (dark green), RTX PRO 6000 Blackwell Max-Q V3 FP8\/V2 (light green), RTX A6000 V3 FP16\/V2 (medium green). RTX PRO 6000 Blackwell Max-Q V3 FP8 reaches the highest speedup on every RTX PRO 6000 Blackwell Max-Q config, peaking at 42\u00d7 on xlarge and 40\u00d7 on wide-c256. RTX A6000 V3 FP16 peaks at the small\/canonical end (~21\u00d7\/19\u00d7) and declines on larger configs where it becomes DRAM-bound.\" class=\"lazyload wp-image-119008\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1.webp 832w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-179x98.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-300x165.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-768x422.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-625x343.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-645x354.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-500x275.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-160x88.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-362x199.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/cross-gpu-speedup-over-v2-fp16-1-200x110.png 200w\" data-sizes=\"(max-width: 832px) 100vw, 832px\"\/><figcaption class=\"wp-element-caption\">Determine 6. Cross-GPU speedup over V2 FP16. RTX A6000 V3 FP16 reaches 11\u201322\u00d7 throughout the examined configurations, whereas RTX PRO 6000 Blackwell Max-Q V3 FP8 reaches 11\u201342\u00d7 relying on level rely and channel width<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"deploy_and_validate_the_tensorrt_plugin\" class=\"wp-block-heading\">Deploy and validate the TensorRT plugin<\/h2>\n<p class=\"wp-block-paragraph\">BEVPoolV3 is uncovered as a TensorRT IPluginV3 operator. The plugin accepts the five-array scatter map plus depth and feat, then dispatches the suitable kernel for the GPU class and dtype. The benchmark path used ONNX-to-TensorRT builds and CUDA Graph replay with trtexec.<\/p>\n<p class=\"wp-block-paragraph\">For validation, evaluate in opposition to an FP64 reference or an current trusted V2 path. The RTX A6000 DRAM-adapted kernel handed all examined output parts throughout the six configurations at atol=1e-2, with most noticed error of 0.0065. On RTX PRO 6000 Blackwell Max-Q, V2 and V3 produced similar outputs for the examined configurations, indicating that the optimized scatter-map and launch adjustments preserved the numerical conduct of the reference path.<\/p>\n<h2 id=\"map_the_algorithm_onto_the_hardware\" class=\"wp-block-heading\">Map the algorithm onto the {hardware}<\/h2>\n<p class=\"wp-block-paragraph\">The 4 BEVPoolV3 algorithmic adjustments are moveable, however the manufacturing kernel should match the lively GPU bottleneck. The important thing resolution is whether or not the BEV-pooling working set matches in L2.<\/p>\n<p class=\"wp-block-paragraph\">On RTX A6000, the canonical working set exceeds L2, so the kernel is restricted by random-gather DRAM visitors. The FP16 path subsequently prioritizes byte discount and cache preservation. Rising TILE_C from 8 to 16 cuts the C=80 tile passes from 10 to five, lowering loop overhead and repeated scalar work. Utilizing __half2 accumulation with __hfma2 avoids pointless FP16-to-FP32 widening and packing. Cache-streaming output shops forestall the 12.8 MB output tensor from evicting the smaller L2-resident index arrays. After these adjustments, the RTX A6000 path reaches 90.0 \u00b5s within the canonical config, in contrast with 1,738.0 \u00b5s for V2 FP16.<\/p>\n<p class=\"wp-block-paragraph\">On RTX PRO 6000 Blackwell Max-Q, the canonical working set matches in L2, so the limiting components shift towards instruction subject, occupancy, and dependency latency. The manufacturing kernel first matches the high-occupancy V2+DO-style launch envelope, then removes inner-loop overhead with the five-array scatter map and precomputed indices. This avoids runtime integer division and reduces scatter-map load strain. Within the canonical config, V3 FP16 reaches 17.3 \u00b5s versus 37.8 \u00b5s for V2+DO FP16, a 2.18x speedup on the similar dtype.<\/p>\n<p class=\"wp-block-paragraph\">The FP8 path additional specializes within the large-L2 case. As a result of function and output information are served from L2, lowering their dtype can translate into actual latency positive aspects. The manufacturing FP8 path makes use of per-channel-count entry factors, LDG.64 index packing for C=80, and wider function masses for C=128 and C=256. Extra aggressive combos, akin to including loop unrolling on high of the packed-index path, didn&#8217;t compose cleanly as a result of they elevated register strain and spill visitors.<\/p>\n<p class=\"wp-block-paragraph\">The precision ladder has a sensible vacation spot, and our NVFP4 analysis helps make clear precisely the place every format shines: we examined an NVFP4 path that shops digicam options in E2M1 with per-16-element E4M3 microblock scales whereas conserving depth and output in FP8, and even with an aggressively optimized implementation that includes __half2 packed accumulators, fused scale\u2013depth coefficients, and a half-precision LUT, the decode overhead causes it to run notably slower than the FP8 baseline.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Profiling with Nsight Compute exhibits the kernel is totally resident in L2 cache, with low DRAM bandwidth utilization and smsp__issue_active hovering properly under peak throughput, whereas the ALU pipeline carries considerably extra directions than the FMA pipeline.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">This means that this scatter-reduce regime has already captured the out there byte-efficiency advantages at FP8, whereas the NVFP4 further per-element nibble extraction, worth decode, and per-microblock scale fold introduce inner-loop work that the FP8 path avoids by way of a single scalar FP8 to half conversion. The result&#8217;s a crisp workload-placement story: NVFP4 stays an extremely highly effective match for compute-bound matrix multiplication shapes flowing by way of Tensor Cores by way of MMA.type::nvfp4, whereas for L2-resident scatter-reduce workloads, FP8 is good on the dtype ladder.<\/p>\n<p class=\"wp-block-paragraph\">The identical evaluation applies past BEV pooling. For sparse embeddings, voxelization, histograms, segmented reductions, and different gather- or scatter-heavy operators, first classify the reminiscence regime, then use Nsight Compute to find out whether or not the lively ceiling is bandwidth, instruction subject, or occupancy.<\/p>\n<p class=\"wp-block-paragraph\">Desk 1 summarizes RTX PRO 6000 Blackwell Max-Q TensorRT plugin-path latency, reported as 100-iteration median latency.<\/p>\n<figure class=\"wp-block-table aligncenter\">ConfigC-DimensionV2 FP16V2+DO FP16V3 FP16V3 FP8V3 FP8 \/ V2small80137.8 \u00b5s31.5 \u00b5s12.7 \u00b5s12.6 \u00b5s10.94xcanonical80274.0 \u00b5s37.8 \u00b5s17.3 \u00b5s16.4 \u00b5s16.71xlarge80749.9 \u00b5s48.0 \u00b5s27.3 \u00b5s24.9 \u00b5s30.12xxlarge801,675.0 \u00b5s61.9 \u00b5s48.0 \u00b5s39.8 \u00b5s42.09xwide_c128128457.3 \u00b5s54.2 \u00b5s21.4 \u00b5s14.8 \u00b5s30.90xwide_c256256880.9 \u00b5s152.3 \u00b5s33.4 \u00b5s22.0 \u00b5s40.04x<figcaption class=\"wp-element-caption\">Desk 1. TensorRT plugin-path latency on RTX PRO 6000 Blackwell Max-Q throughout a number of mannequin configurations. Values report 100-iteration median latency per benchmarked inference\/plugin name in microseconds for V2 FP16, V2+DO FP16, V3 FP16, and V3 FP8. The ultimate column exhibits the speedup of V3 FP8 relative to V2 FP16, computed from latency discount<\/figcaption><\/figure>\n<h2 id=\"considerations_for_edge-class_platforms\" class=\"wp-block-heading\">Concerns for edge-class platforms<\/h2>\n<p class=\"wp-block-paragraph\">The identical evaluation can prolong to edge-class NVIDIA platforms, together with NVIDIA DRIVE AGX Thor. In early edge-oriented experiments, the FP16 BEVPoolV3 path carries over properly as a result of the core enhancements\u2014eradicating redundant scatter visitors, avoiding runtime index decoding, and utilizing interval-owned writes\u2014are architecture-independent.<\/p>\n<p class=\"wp-block-paragraph\">FP8 speedup, nonetheless, is just not automated. On edge-class targets, smaller drawback sizes, reminiscence hierarchy conduct, register strain, and FP8 conversion overhead can restrict or offset the theoretical dtype bandwidth profit. This makes FP8 a kernel- and architecture-specific optimization fairly than a assured drop-in alternative for FP16.<\/p>\n<h2 id=\"get_started_with_bev_pooling_optimization\" class=\"wp-block-heading\">Get began with BEV pooling optimization<\/h2>\n<p class=\"wp-block-paragraph\">To use the BEVPoolV3 workflow to your personal BEV notion or collect\/scatter-heavy workload, begin by profiling the operator in isolation. Measure the function, depth, scatter-index, and output tensor sizes, then evaluate the entire working set with the goal GPU L2 cache capability.<\/p>\n<p class=\"wp-block-paragraph\">Use NVIDIA Nsight Compute to validate whether or not the lively bottleneck is reminiscence bandwidth, instruction subject, occupancy, or dependency latency. Then select the optimization technique that matches the reminiscence regime: byte discount and cache-preserving shops for DRAM-bound workloads, or occupancy, precomputed indices, vectorized masses, and dtype specialization for L2-resident workloads.<\/p>\n<p class=\"wp-block-paragraph\">The identical strategy applies to sparse embeddings, voxelization, histograms, segmented reductions, and different irregular memory-bound kernels. Use the BEVPoolV3 outcomes as a information for profiling your personal operator, choosing the precise implementation technique for the goal GPU, and validating the outcome earlier than deploying by way of TensorRT. For associated assets, see the TensorRT documentation, CUDA C++ Programming Information, Nsight Compute documentation, and NVIDIA Developer Boards.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>An more and more widespread design sample for autonomous autos (AVs), robotics, and spatial AI techniques is fowl\u2019s-eye-view (BEV) notion. BEV fashions venture multicamera picture options right into a shared top-down grid, offering downstream notion and planning modules with a standard spatial format for reasoning about lanes, autos, pedestrians, and free house. A key operation [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1466,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/tensorrt-optimized-industries-1.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[1913,1666,1931,1597,81,953,1932],"class_list":["post-1464","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-accelerating","tag-applications","tag-bev","tag-gpus","tag-nvidia","tag-physical","tag-pooling"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes - Future News 24<\/title>\n<meta name=\"description\" content=\"An increasingly common design pattern for autonomous vehicles (AVs), robotics, and spatial AI systems is bird&rsquo;s&#x2d;eye&#x2d;view (BEV) perception.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes - Future News 24\" \/>\n<meta property=\"og:description\" content=\"An increasingly common design pattern for autonomous vehicles (AVs), robotics, and spatial AI systems is bird&rsquo;s&#x2d;eye&#x2d;view (BEV) perception.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-24T16:30:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-25T09:59:25+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/tensorrt-optimized-industries-1.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/tensorrt-optimized-industries-1.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"14 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes\",\"datePublished\":\"2026-06-24T16:30:00+00:00\",\"dateModified\":\"2026-06-25T09:59:25+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/\"},\"wordCount\":2907,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/tensorrt-optimized-industries-1.webp\",\"keywords\":[\"Accelerating\",\"applications\",\"BEV\",\"GPUs\",\"NVIDIA\",\"Physical\",\"Pooling\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/\",\"name\":\"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/tensorrt-optimized-industries-1.webp\",\"datePublished\":\"2026-06-24T16:30:00+00:00\",\"dateModified\":\"2026-06-25T09:59:25+00:00\",\"description\":\"An increasingly common design pattern for autonomous vehicles (AVs), robotics, and spatial AI systems is bird&rsquo;s&#x2d;eye&#x2d;view (BEV) perception.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/tensorrt-optimized-industries-1.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/tensorrt-optimized-industries-1.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/24\\\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes - Future News 24","description":"An increasingly common design pattern for autonomous vehicles (AVs), robotics, and spatial AI systems is bird&rsquo;s&#x2d;eye&#x2d;view (BEV) perception.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/","og_locale":"en_US","og_type":"article","og_title":"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes - Future News 24","og_description":"An increasingly common design pattern for autonomous vehicles (AVs), robotics, and spatial AI systems is bird&rsquo;s&#x2d;eye&#x2d;view (BEV) perception.","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/","og_site_name":"Future News 24","article_published_time":"2026-06-24T16:30:00+00:00","article_modified_time":"2026-06-25T09:59:25+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/tensorrt-optimized-industries-1.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/tensorrt-optimized-industries-1.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"14 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes","datePublished":"2026-06-24T16:30:00+00:00","dateModified":"2026-06-25T09:59:25+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/"},"wordCount":2907,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/tensorrt-optimized-industries-1.webp","keywords":["Accelerating","applications","BEV","GPUs","NVIDIA","Physical","Pooling"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/","name":"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/tensorrt-optimized-industries-1.webp","datePublished":"2026-06-24T16:30:00+00:00","dateModified":"2026-06-25T09:59:25+00:00","description":"An increasingly common design pattern for autonomous vehicles (AVs), robotics, and spatial AI systems is bird&rsquo;s&#x2d;eye&#x2d;view (BEV) perception.","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/tensorrt-optimized-industries-1.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/tensorrt-optimized-industries-1.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/24\/accelerating-bev-pooling-on-nvidia-gpus-for-physical-ai-applications\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Accelerating BEV Pooling on NVIDIA GPUs for Bodily AI Purposes"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1464","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=1464"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1464\/revisions"}],"predecessor-version":[{"id":1465,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1464\/revisions\/1465"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/1466"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=1464"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=1464"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=1464"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}