{"id":1715,"date":"2026-06-30T17:36:00","date_gmt":"2026-06-30T17:36:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/"},"modified":"2026-07-01T05:59:07","modified_gmt":"2026-07-01T05:59:07","slug":"designing-gpu-accelerated-query-engines-with-nvidia-gqe","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/","title":{"rendered":"Designing GPU-Accelerated Question Engines with NVIDIA GQE"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">GPU-accelerated question engines are sometimes constrained by reminiscence and I\/O bandwidth. NVIDIA {hardware} advances\u2014together with excessive bandwidth reminiscence (HBM), NVIDIA NVLink-C2C, and devoted decompression engines featured in NVIDIA GB200 NVL4\u2014assist take away these bottlenecks by rising efficient storage capability, accelerating knowledge motion between CPUs and GPUs, and rushing knowledge entry with out consuming streaming multiprocessor (SM) sources.<\/p>\n<p class=\"wp-block-paragraph\">On this publish, we present how databases can use these applied sciences to speed up GPU question execution. You\u2019ll study methods for environment friendly CPU-GPU knowledge motion, compression, partition pruning, and overlapping knowledge switch with computation.<\/p>\n<h2 id=\"architecture_overview_of_gqe\" class=\"wp-block-heading\">Structure overview of GQE<\/h2>\n<p class=\"wp-block-paragraph\">GQE (GPU Question Engine) is a reference structure designed to execute SQL queries at excessive efficiency over giant knowledge units on fashionable NVIDIA {hardware}. Underneath the hood, GQE makes use of NVIDIA cuDF and different NVIDIA CUDA-X libraries, together with CCCL, nvCOMP, and nvSHMEM.<\/p>\n<p class=\"wp-block-paragraph\">GQE can assist affect question engines to:<\/p>\n<p>Transfer execution to GPUs.<\/p>\n<p>Transfer decompression to nvCOMP.<\/p>\n<p>Make knowledge codecs GPU-friendly.<\/p>\n<p>Shut end-to-end efficiency gaps when working on GPUs.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a44aca8d623e&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a44aca8d623e\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1920\" height=\"1080\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1.webp\" alt=\"Architecture diagram tracing a SQL query from parsing and Substrait plan optimization, through physical plan generation and task graph construction, down to GPU-accelerated reads, joins and sorts, built on cuDF and nvCOMP.\" class=\"wp-image-119306\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1.webp 1920w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-179x101.jpg 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-300x169.jpg 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-768x432.jpg 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-625x352.jpg 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-1536x864.jpg 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-645x363.jpg 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-660x370.jpg 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-500x281.jpg 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-160x90.jpg 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-362x204.jpg 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-196x110.jpg 196w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-1024x576.jpg 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-960x540.jpg 960w\" sizes=\"(max-width: 1920px) 100vw, 1920px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1920\" height=\"1080\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1.webp\" alt=\"Architecture diagram tracing a SQL query from parsing and Substrait plan optimization, through physical plan generation and task graph construction, down to GPU-accelerated reads, joins and sorts, built on cuDF and nvCOMP.\" class=\"lazyload wp-image-119306\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1.webp 1920w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-179x101.jpg 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-300x169.jpg 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-768x432.jpg 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-625x352.jpg 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-1536x864.jpg 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-645x363.jpg 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-660x370.jpg 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-500x281.jpg 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-160x90.jpg 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-362x204.jpg 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-196x110.jpg 196w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-1024x576.jpg 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/agentic-ai-diagrams-architecture-overview-gqe-fig-2-1920x1080-1-960x540.jpg 960w\" data-sizes=\"(max-width: 1920px) 100vw, 1920px\"\/><figcaption class=\"wp-element-caption\">Determine 1. A SQL question flows by GQE\u2019s three structure layers\u2014question, knowledge, and execution\u2014to turn into GPU-accelerated<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">In Determine 1, we give an outline of the system design by breaking down GQE into a question,\u00a0 knowledge, and execution layer. These handle the transition from a SQL question and enter knowledge to hardware-level execution. The layers match collectively as follows.<\/p>\n<p class=\"wp-block-paragraph\">The question layer enhances the execution engine with a SQL parser and a question optimizer. The question layer natively accepts Substrait plans, an open-source question plan format, for execution in GQE. Substrait makes it doable to judge the advantages of GPU execution by exporting question plans from an present database product and working the plan in GQE. In Determine 2, Apache DataFusion transforms a SQL string right into a Substrait plan. GQE consumes that plan as an optimized logical question plan, provides GQE-specific refinements, and transforms the question right into a bodily plan.<\/p>\n<p class=\"wp-block-paragraph\">The info layer shops and organizes person knowledge for quick entry by the executor. In GQE, storage is abstracted into pluggable, specialised readers that deal with totally different knowledge codecs and storage mediums\u2014it at present helps GPU reminiscence, CPU reminiscence, and disk. On this publish, we deal with the high-performance GQE in-memory desk format and assume this knowledge is saved in CPU reminiscence. GQE transfers knowledge chunks to the GPU on-demand to saturate the GPU with work with out storing the complete dataset in GPU reminiscence. When a piece arrives on the GPU, the information layer palms off to the execution layer.<\/p>\n<p class=\"wp-block-paragraph\">The execution layer executes the bodily question plan in opposition to the information to provide question outcomes. GQE generates the bodily plan right into a job graph, which defines the execution schedule. The duty graph accommodates relational operators constructed on the open-source NVIDIA cuDF library, which implements the operators in extremely optimized CUDA C++ code. As a result of the information layer transfers in chunks, GQE can decompose operators and execute duties on these chunks concurrently as pipelined CUDA streams.<\/p>\n<p class=\"wp-block-paragraph\">In abstract, GQE unlocks the excessive throughput of the {hardware} by a GPU-native design.<\/p>\n<h2 id=\"data_layout_and_transfer_orchestration\" class=\"wp-block-heading\">Information structure and switch orchestration<\/h2>\n<p class=\"wp-block-paragraph\">The GQE knowledge layer is optimized to effectively switch knowledge from host reminiscence to machine reminiscence. We reduce knowledge switch latency by maximizing throughput and lowering the quantity of information moved. Within the following, we give an outline of our in-memory knowledge structure and the host-to-device switch orchestration, that are instrumental to minimizing switch latency.<\/p>\n<p class=\"wp-block-paragraph\">GQE design objectives\u00a0<\/p>\n<p class=\"wp-block-paragraph\">As GQE builds on cuDF, the design assumes that in-GPU knowledge is structured as cuDF-native tables. Nonetheless, the host reminiscence structure can optimize transfers for NVIDIA NVLink C2C and PCIe. cudaMemcpy is the usual switch methodology. On this method, the CPU orchestrates GPU execution and copies knowledge in a bulk switch. This additionally varieties the premise for compressed transfers.<\/p>\n<p class=\"wp-block-paragraph\">Information structure<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a44aca8d704d&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a44aca8d704d\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"1504\" height=\"1120\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26.webp\" alt=\"Diagram of an in-memory table divided into row groups, each containing metadata and columnar partitions, with an arrow showing conversion into a cuDF table on the GPU.\" class=\"wp-image-119243\" style=\"aspect-ratio:1.342892114005945;width:488px;height:auto\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26.webp 1504w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-154x115.png 154w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-300x223.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-768x572.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-625x465.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-645x480.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-403x300.png 403w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-121x90.png 121w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-362x270.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-148x110.png 148w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-1024x763.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-725x540.png 725w\" sizes=\"(max-width: 1504px) 100vw, 1504px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1504\" height=\"1120\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26.webp\" alt=\"Diagram of an in-memory table divided into row groups, each containing metadata and columnar partitions, with an arrow showing conversion into a cuDF table on the GPU.\" class=\"lazyload wp-image-119243\" style=\"aspect-ratio:1.342892114005945;width:488px;height:auto\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26.webp 1504w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-154x115.png 154w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-300x223.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-768x572.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-625x465.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-645x480.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-403x300.png 403w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-121x90.png 121w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-362x270.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-148x110.png 148w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-1024x763.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-26-725x540.png 725w\" data-sizes=\"(max-width: 1504px) 100vw, 1504px\"\/><figcaption class=\"wp-element-caption\">Determine 2. GQE\u2019s in-memory desk format organizes columnar knowledge into row teams and partitions for environment friendly switch to the GPU<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Determine 2 exhibits the desk knowledge structure, which is horizontally subdivided into row teams. Every row group consists of columns and encapsulates metadata. Inside a row group, GQE shops columns as non-contiguous partitions. Throughout a switch, the storage layer converts a set of partitions right into a cuDF column. Thus, the information layer hides the implementation particulars of compression and partition pruning from the execution layer.<\/p>\n<p class=\"wp-block-paragraph\">Switch Orchestration<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a44aca8d7a3b&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a44aca8d7a3b\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"785\" height=\"306\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29.webp\" alt=\"Timeline showing four row groups moving through overlapping pipeline stages, including host scheduling, H2D transfer, decompression, and CUDA kernel execution across concurrent CUDA streams.\u00a0\" class=\"wp-image-119248\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29.webp 785w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-179x70.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-300x117.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-768x299.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-625x244.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-645x251.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-500x195.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-160x62.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-362x141.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-282x110.png 282w\" sizes=\"(max-width: 785px) 100vw, 785px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"785\" height=\"306\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29.webp\" alt=\"Timeline showing four row groups moving through overlapping pipeline stages, including host scheduling, H2D transfer, decompression, and CUDA kernel execution across concurrent CUDA streams.\u00a0\" class=\"lazyload wp-image-119248\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29.webp 785w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-179x70.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-300x117.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-768x299.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-625x244.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-645x251.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-500x195.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-160x62.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-362x141.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-29-282x110.png 282w\" data-sizes=\"(max-width: 785px) 100vw, 785px\"\/><figcaption class=\"wp-element-caption\">Determine 3. Pipeline parallelism overlaps scheduling, knowledge switch, decompression, and GPU execution throughout row teams to hurry up transfers<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">In Determine 3, we present how the CPU orchestrates a switch. Following finest CUDA apply, transfers use pipeline parallelism to effectively make the most of {hardware} parts. A pipelined switch consists of a number of phases. In compressed, partitioned knowledge, there are 4 phases.<\/p>\n<p>In Stage 0, a number thread performs scheduling. Scheduling entails computing the reminiscence vary to switch, allocating a vacation spot buffer, and invoking the required CUDA strategies.<\/p>\n<p>In Stage 1, the GPU performs the H2D switch.<\/p>\n<p>Stage 2 decompresses the information.<\/p>\n<p>Stage 3, added outdoors the information layer, wherein the CUDA kernels compute the question.<\/p>\n<p class=\"wp-block-paragraph\">These 4 phases ought to overlap. Ideally, the question runtime equals the longest-running stage, and all remaining phases are hidden by the pipeline.<\/p>\n<h2 id=\"data_transfer_optimizations\" class=\"wp-block-heading\">Information switch optimizations<\/h2>\n<p class=\"wp-block-paragraph\">Quick knowledge entry performs a major position within the efficiency benefit achieved by GQE. The primary knowledge entry optimizations employed are compression and partition pruning. Within the following, we describe how these optimizations work.<\/p>\n<h3 id=\"compression\" class=\"wp-block-heading\">Compression<\/h3>\n<p class=\"wp-block-paragraph\">GQE receives two predominant advantages from compression: question dataset capability and question acceleration. Compression allows a question engine to develop the dataset dimension that may be processed utilizing a given reminiscence allotment by lowering the general in-memory footprint. Information switch of compressed buffers, mixed with quick decompression by the GPU, hastens transfers even on quick interconnects like NVLink C2C. GQE compresses the datasets with GPU-optimized codecs that enhance compression ratios and supply superior GPU decompression speeds in comparison with utilizing legacy codecs.<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA nvCOMP library\u00a0<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA nvCOMP is a library for GPU-accelerated compression and decompression. It gives a spread of normal and GPU-optimized compression codecs. The person can decide from the supported algorithms to stability compression ratio, compression, and decompression throughput. nvCOMP can wrap CPU libraries equivalent to lz4hc inside its high-level interface, offering extra configuration choices. GQE makes use of nvCOMP for its compression and decompression routines.<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA Blackwell Decompression Engine<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA launched a brand new Decompression Engine (DE) within the NVIDIA Blackwell structure that allows nvCOMP to rapidly decompress LZ77-based codecs like LZ4, Snappy, and Deflate with out utilizing SM sources. Decompression with DE, SM kernels, and CE copies can totally overlap when utilizing a number of CUDA streams.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">DE on a single NVIDIA Blackwell B200 GPU can attain as much as 400 GB\/s in database purposes. For instance, at a 4x compression ratio, it achieves roughly 400 GB\/s efficient host-to-device throughput whereas leaving 100 GB\/s C2C host-to-device bandwidth obtainable. The remaining bandwidth will be harnessed to switch different knowledge, together with encoded knowledge that&#8217;s decompressed on the SMs.<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA GQE\u2019s compression method\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Determine 4 exhibits the hybrid compression method, which makes use of light-weight algorithms, equivalent to Cascaded, to make use of particular patterns within the structured knowledge the place doable, and the DE when LZ-based algorithms are wanted to attain good compression ratios.<\/p>\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a44aca8d8c5a&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a44aca8d8c5a\" class=\"wp-block-image size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1400\" height=\"780\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30.webp\" alt=\"Four-panel walkthrough of cascaded compression: an original column of 16-bit values is reduced by delta encoding, then run-length encoding, and finally bit-packing into a compact binary representation.\" class=\"wp-image-119249\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30.webp 1400w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-179x100.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-300x167.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-768x428.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-625x348.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-645x359.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-500x279.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-160x90.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-362x202.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-197x110.png 197w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-1024x571.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-960x535.png 960w\" sizes=\"(max-width: 1400px) 100vw, 1400px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1400\" height=\"780\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30.webp\" alt=\"Four-panel walkthrough of cascaded compression: an original column of 16-bit values is reduced by delta encoding, then run-length encoding, and finally bit-packing into a compact binary representation.\" class=\"lazyload wp-image-119249\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30.webp 1400w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-179x100.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-300x167.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-768x428.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-625x348.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-645x359.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-500x279.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-160x90.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-362x202.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-197x110.png 197w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-1024x571.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-30-960x535.png 960w\" data-sizes=\"(max-width: 1400px) 100vw, 1400px\"\/><figcaption class=\"wp-element-caption\">Determine 4. nvCOMP Cascaded format chains delta encoding, run-length encoding, and bit-packing to effectively compress columnar knowledge<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">When contemplating tips on how to compress a given column, a question engine has a couple of choices. It might require customers to specify an algorithm for every column, however that is unwieldy for very giant databases. The method we\u2019ve taken is to try each LZ4 and Cascaded. LZ4 is our alternative for generic knowledge as a result of it achieves excessive ratios in comparison with different LZ77-only compressors, and is supported by the Decompression Engine.<\/p>\n<p class=\"wp-block-paragraph\">To find out the compression algorithm to make use of, we compress the information utilizing each LZ4 and Cascaded algorithms. Cascaded can obtain extraordinarily quick compression charges, at roughly 500 GB\/s on B200. This permits us to strive the additional algorithm with out important overhead within the knowledge loading stage.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">We stability when to make use of Cascaded vs LZ4 utilizing two heuristics:\u00a0<\/p>\n<p>Cascaded and LZ4 have totally different compression ratio thresholds, which set up minimums for us to make use of that algorithm.<\/p>\n<p>Cascaded should obtain the next compression ratio than LZ4 to be chosen over LZ4. The set off is a configurable a number of of the LZ4 compression ratio.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">We use the selection of algorithm to assist stability C2C bandwidth, DE, and SM sources.<\/p>\n<h2 id=\"partition_pruning\" class=\"wp-block-heading\">Partition pruning<\/h2>\n<p class=\"wp-block-paragraph\">Earlier than transferring knowledge from the CPU to the GPU, GQE employs filter pruning to skip partitions that don\u2019t contribute to the question end result. This mechanism depends on metadata summarizing the desk contents and the predicates outlined within the SQL question.<\/p>\n<p class=\"wp-block-paragraph\">Metadata and storage<\/p>\n<p class=\"wp-block-paragraph\">GQE makes use of zone maps to assist filter pruning. When knowledge is loaded as in-memory tables, GQE horizontally splits the desk into row teams and fixed-size partitions with a default of 10M rows. For every partition, GQE computes the minimal and most values for each column and shops this metadata as cuDF tables in GPU reminiscence, so pruning can run with out changing into a bottleneck. Computing the zone maps provides about 1% to the preliminary Parquet load time and occurs solely as soon as, not throughout question execution.<\/p>\n<p class=\"wp-block-paragraph\">Pruning and job orchestration<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a44aca8d9bb4&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a44aca8d9bb4\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1886\" height=\"1290\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27.webp\" alt=\"Diagram showing a query predicate compared against min\/max zone-map metadata for each partition. Partitions that cannot match are pruned; only retained partitions are transferred and assembled into a cuDF table on the GPU.\" class=\"wp-image-119244\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27.webp 1886w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-168x115.png 168w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-300x205.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-768x525.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-625x427.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-1536x1051.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-645x441.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-439x300.png 439w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-132x90.png 132w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-362x248.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-161x110.png 161w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-1024x700.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-789x540.png 789w\" sizes=\"(max-width: 1886px) 100vw, 1886px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1886\" height=\"1290\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27.webp\" alt=\"Diagram showing a query predicate compared against min\/max zone-map metadata for each partition. Partitions that cannot match are pruned; only retained partitions are transferred and assembled into a cuDF table on the GPU.\" class=\"lazyload wp-image-119244\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27.webp 1886w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-168x115.png 168w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-300x205.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-768x525.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-625x427.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-1536x1051.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-645x441.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-439x300.png 439w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-132x90.png 132w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-362x248.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-161x110.png 161w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-1024x700.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-27-789x540.png 789w\" data-sizes=\"(max-width: 1886px) 100vw, 1886px\"\/><figcaption class=\"wp-element-caption\">Determine 5. Filter pruning evaluates question predicates in opposition to partition-level zone maps to skip irrelevant knowledge earlier than switch, lowering knowledge motion to the GPU<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Determine 5 exhibits the filter pruning course of. Throughout job graph development, GQE derives a pruning expression by reworking question predicates into comparisons in opposition to the row teams\u2019 zone maps. Partitions that may\u2019t contribute to the question end result are pruned. On this instance, partition 1 is pruned as a result of the zone map signifies that every one values saved on this partition are lower than 9, and due to this fact additionally lower than 15, the decrease sure. The remaining partitions are transferred to GPU reminiscence and decompressed if essential. Even when partitions are discontiguous in CPU reminiscence after pruning, e.g., as a result of they&#8217;re contained in a number of row teams, they&#8217;re transferred and assembled right into a contiguous reminiscence block wrapped as a cuDF desk on the GPU.<\/p>\n<p class=\"wp-block-paragraph\">Filter pruning in GQE is extremely efficient. Within the TPC-H benchmark utilizing the 1 TB scale dataset, filter pruning skips 31% of information throughout all 22 queries. The influence is an end-to-end speedup of 1.43\u00d7.<\/p>\n<p class=\"wp-block-paragraph\">The analysis of zone maps provides minimal overhead, on common, 2.2 ms for benchmark queries on 1 TB of information.<\/p>\n<p class=\"wp-block-paragraph\">Information switch optimizations<\/p>\n<p class=\"wp-block-paragraph\">In GQE, we conceive a novel batched switch optimization for partitions.<\/p>\n<p class=\"wp-block-paragraph\">A number of partitions are transferred to the GPU in a single batch utilizing cudaMemcpyBatchAsync, lowering overhead for fine-grained partitions. Batching additionally helps keep away from delays from interleaved CUDA streams. When partitions are transferred individually, transfers from different streams can delay the subsequent kernel launch. Transferring partitions in the identical batch avoids this delay.<\/p>\n<h2 id=\"performance_highlights\" class=\"wp-block-heading\">Efficiency highlights<\/h2>\n<p class=\"wp-block-paragraph\">To judge the B200 GPU options mentioned above in a full Grace Blackwell system, we benchmarked GQE on TPC-H at Scale Issue 1000 (1TB) utilizing one of many two B200 GPUs in an NVIDIA GB200 NVL4 server, the place B200 GPUs are linked to the Grace CPU with NVLink-C2C.\u00a0 We used DuckDB 1.4.1 on the Turin Epyc 9755 CPU because the baseline. Every question was averaged over 5 hot-cache runs, with compression and pruning enabled on each side. We tune the GQE parameters per question, together with the diploma of parallelism and bodily operator planning.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The TPC-H dataset is optimized for partition pruning and compression by clustering the lineitem desk on l_shipdate and the orders desk on o_orderdate, and partitioning each tables by month. Internally, every partition is sorted on l_orderkey and o_orderkey, respectively.<\/p>\n<p class=\"wp-block-paragraph\">In Determine 6, we present the runtime of the 22 queries. GQE outperforms DuckDB on 20 of twenty-two queries, with the most important beneficial properties on Q11, Q14, and Q15, the place partition pruning and compression sharply reduce knowledge motion throughout NVLink C2C. GQE showcases that even bandwidth-heavy queries like Q1 and Q6 execute rapidly on the GPU with these optimizations. In sum, GQE runs all queries in 9.0 s, in comparison with 74.0 s and 70.6 s for DuckDB in single and dual-socket configurations, respectively.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a44aca8dae93&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a44aca8dae93\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1200\" height=\"600\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31.webp\" alt=\"Bar chart comparison for queries-per-second with GQE on NVIDIA GB200 versus DuckDB on one- and two-socket AMD Turin EPYC 9755 across TPC-H SF1000 queries Q1\u2013Q22, with GQE achieving higher throughput on over 90% of queries based on TPC-H.\" class=\"wp-image-119252\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31.webp 1200w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-179x90.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-300x150.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-768x384.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-625x313.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-645x323.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-500x250.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-160x80.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-362x181.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-220x110.png 220w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-1024x512.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-960x480.png 960w\" sizes=\"(max-width: 1200px) 100vw, 1200px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1200\" height=\"600\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31.webp\" alt=\"Bar chart comparison for queries-per-second with GQE on NVIDIA GB200 versus DuckDB on one- and two-socket AMD Turin EPYC 9755 across TPC-H SF1000 queries Q1\u2013Q22, with GQE achieving higher throughput on over 90% of queries based on TPC-H.\" class=\"lazyload wp-image-119252\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31.webp 1200w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-179x90.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-300x150.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-768x384.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-625x313.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-645x323.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-500x250.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-160x80.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-362x181.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-220x110.png 220w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-1024x512.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-31-960x480.png 960w\" data-sizes=\"(max-width: 1200px) 100vw, 1200px\"\/><figcaption class=\"wp-element-caption\">Determine 6. Outcomes from NVIDIA testing, GQE on a single GB200 GPU outperforms DuckDB on dual-socket AMD Turin CPUs in over 90% of queries based mostly on TPC-H at 1 TB scale issue<\/figcaption><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a44aca8db77e&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a44aca8db77e\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1200\" height=\"600\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32.webp\" alt=\"Bar chart showing the performance speedup of GQE on NVIDIA GB200 compared to the best CPU configuration for TPC-H SF1000 queries, illustrating an aggregate 7.5x speedup with per-query gains ranging from near parity to over 25x.\" class=\"wp-image-119254\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32.webp 1200w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-179x90.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-300x150.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-768x384.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-625x313.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-645x323.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-500x250.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-160x80.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-362x181.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-220x110.png 220w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-1024x512.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-960x480.png 960w\" sizes=\"(max-width: 1200px) 100vw, 1200px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1200\" height=\"600\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32.webp\" alt=\"Bar chart showing the performance speedup of GQE on NVIDIA GB200 compared to the best CPU configuration for TPC-H SF1000 queries, illustrating an aggregate 7.5x speedup with per-query gains ranging from near parity to over 25x.\" class=\"lazyload wp-image-119254\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32.webp 1200w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-179x90.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-300x150.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-768x384.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-625x313.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-645x323.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-500x250.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-160x80.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-362x181.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-220x110.png 220w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-1024x512.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image-32-960x480.png 960w\" data-sizes=\"(max-width: 1200px) 100vw, 1200px\"\/><figcaption class=\"wp-element-caption\">Determine 7. GQE delivers a 7.5x mixture speedup over one of the best CPU configuration, with per-query beneficial properties starting from close to parity to over 25x<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">We current the speedups in Determine 7. GQE delivers as much as 25.5x over DuckDB\u2019s finest CPU socket configuration, outperforming it on 20 of twenty-two queries and reaching 3x or larger on 17. Aggregated throughout all queries, GQE on the GB200 achieves a 7.5x speedup on complete execution time.<\/p>\n<p class=\"wp-block-paragraph\">The check outcomes on this weblog publish are derived from TPC-H choice assist benchmark and aren\u2019t corresponding to printed TPC-H outcomes, because the check outcomes on this weblog don&#8217;t adjust to the TPC-H specification.<\/p>\n<h2 id=\"apply_gqe_best_practices_to_data_platforms\" class=\"wp-block-heading\">Apply GQE finest practices to knowledge platforms<\/h2>\n<p class=\"wp-block-paragraph\">Database engines can translate NVIDIA Grace Blackwell {hardware} options into measurable question efficiency beneficial properties with focused optimizations. In GQE, partition pruning and hybrid compression reduce switch quantity whereas NVLink-C2C and DE {hardware} improve switch throughput. These optimizations cut back switch time and compose into subtle question execution utilizing NVIDIA cuDF, NVIDIA nvCOMP, and different CUDA-X libraries.<\/p>\n<p class=\"wp-block-paragraph\">On TPC-H SF1000, GQE achieved a 7.5x speedup on complete execution time over a state-of-the-art CPU database, exhibiting how knowledge structure, compression technique, and execution will be designed collectively for contemporary database engines.<\/p>\n<p class=\"wp-block-paragraph\">Leverage the GQE open-source reference structure and design, and efficiency optimizations, and discover how GQE can speed up your knowledge platforms.\u00a0<\/p>\n<h3 id=\"acknowledgements\" class=\"wp-block-heading\">Acknowledgements<\/h3>\n<p class=\"wp-block-paragraph\">The authors want to thank Tanmay Gujar for his technical contributions to GQE and his assessment of this publish. We additionally prolong our because of all GQE contributors\u2014Hao Gao, Yadu Kiran, James Xia, Eyal Soha, Lingyan Yin, Daniel Juenger, Siyuan Lin, Bret Alfieri, Nico Iskos, Zhengru Wang, Rui Bao, Dhruv Sundararaman, Jiachun Li, and Kate Cheng\u2014for his or her technical contributions. Lastly, we\u2019d prefer to thank Nikolay Sakharnykh and Nuttiiya Seekhao for his or her assessment.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>GPU-accelerated question engines are sometimes constrained by reminiscence and I\/O bandwidth. NVIDIA {hardware} advances\u2014together with excessive bandwidth reminiscence (HBM), NVIDIA NVLink-C2C, and devoted decompression engines featured in NVIDIA GB200 NVL4\u2014assist take away these bottlenecks by rising efficient storage capability, accelerating knowledge motion between CPUs and GPUs, and rushing knowledge entry with out consuming streaming multiprocessor [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1717,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/GQE.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[2196,2199,2197,2200,81,2198],"class_list":["post-1715","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-designing","tag-engines","tag-gpuaccelerated","tag-gqe","tag-nvidia","tag-query"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Designing GPU-Accelerated Question Engines with NVIDIA GQE - Future News 24<\/title>\n<meta name=\"description\" content=\"GPU&#x2d;accelerated query engines are often constrained by memory and I\/O bandwidth. NVIDIA hardware advances&mdash;including high bandwidth memory (HBM)&#8230;\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Designing GPU-Accelerated Question Engines with NVIDIA GQE - Future News 24\" \/>\n<meta property=\"og:description\" content=\"GPU&#x2d;accelerated query engines are often constrained by memory and I\/O bandwidth. NVIDIA hardware advances&mdash;including high bandwidth memory (HBM)&#8230;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-30T17:36:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-01T05:59:07+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/GQE.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/GQE.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"12 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Designing GPU-Accelerated Question Engines with NVIDIA GQE\",\"datePublished\":\"2026-06-30T17:36:00+00:00\",\"dateModified\":\"2026-07-01T05:59:07+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/\"},\"wordCount\":2426,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/GQE.webp\",\"keywords\":[\"Designing\",\"Engines\",\"GPUAccelerated\",\"GQE\",\"NVIDIA\",\"Query\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/\",\"name\":\"Designing GPU-Accelerated Question Engines with NVIDIA GQE - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/GQE.webp\",\"datePublished\":\"2026-06-30T17:36:00+00:00\",\"dateModified\":\"2026-07-01T05:59:07+00:00\",\"description\":\"GPU&#x2d;accelerated query engines are often constrained by memory and I\\\/O bandwidth. NVIDIA hardware advances&mdash;including high bandwidth memory (HBM)&#8230;\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/GQE.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/GQE.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/30\\\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Designing GPU-Accelerated Question Engines with NVIDIA GQE\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Designing GPU-Accelerated Question Engines with NVIDIA GQE - Future News 24","description":"GPU&#x2d;accelerated query engines are often constrained by memory and I\/O bandwidth. NVIDIA hardware advances&mdash;including high bandwidth memory (HBM)&#8230;","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/","og_locale":"en_US","og_type":"article","og_title":"Designing GPU-Accelerated Question Engines with NVIDIA GQE - Future News 24","og_description":"GPU&#x2d;accelerated query engines are often constrained by memory and I\/O bandwidth. NVIDIA hardware advances&mdash;including high bandwidth memory (HBM)&#8230;","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/","og_site_name":"Future News 24","article_published_time":"2026-06-30T17:36:00+00:00","article_modified_time":"2026-07-01T05:59:07+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/GQE.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/GQE.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"12 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Designing GPU-Accelerated Question Engines with NVIDIA GQE","datePublished":"2026-06-30T17:36:00+00:00","dateModified":"2026-07-01T05:59:07+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/"},"wordCount":2426,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/GQE.webp","keywords":["Designing","Engines","GPUAccelerated","GQE","NVIDIA","Query"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/","name":"Designing GPU-Accelerated Question Engines with NVIDIA GQE - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/GQE.webp","datePublished":"2026-06-30T17:36:00+00:00","dateModified":"2026-07-01T05:59:07+00:00","description":"GPU&#x2d;accelerated query engines are often constrained by memory and I\/O bandwidth. NVIDIA hardware advances&mdash;including high bandwidth memory (HBM)&#8230;","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/GQE.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/GQE.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/30\/designing-gpu-accelerated-query-engines-with-nvidia-gqe\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Designing GPU-Accelerated Question Engines with NVIDIA GQE"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1715","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=1715"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1715\/revisions"}],"predecessor-version":[{"id":1716,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1715\/revisions\/1716"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/1717"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=1715"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=1715"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=1715"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}