Wednesday, September 16, 2026
No Result
View All Result
Future News 24
Advertisement
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
No Result
View All Result
Future News 24
No Result
View All Result
Home AI Platforms & Apps

Working Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72

Future News 24 by Future News 24
July 9, 2026
in AI Platforms & Apps
0 0
0
Working Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


Presto is an open supply, distributed SQL engine for operating quick, interactive queries on very giant datasets. On NVIDIA GPUs, Presto delivers peak efficiency for analytical question workloads and supplies low latency for customers and brokers. GPU-accelerated Presto brings low latency to your analytical workloads, conserving you and your brokers unblocked and iterating as quick as potential. 

This publish demonstrates environment friendly multi-GPU Presto execution on scaled analytical benchmarks utilizing each single-node NVIDIA DGX B200 and multinode NVIDIA GB200 NVL72. We additionally spotlight the significance of NVIDIA GPUDirect Storage (GDS) for top I/O throughput on NVIDIA GB200 NVL72 when paired with an IBM Storage Scale System knowledge supply.

For extra particulars about GPU-accelerated Presto and the significance of UcxExchange for top efficiency communications between GPU employees, see Accelerating Giant-Scale Information Analytics with GPU-Native Velox and NVIDIA cuDF.

How does GPU-accelerated Presto ship peak efficiency?

GPU-accelerated Presto makes use of NVIDIA cuDF algorithms for peak efficiency and NVIDIA NVLink for the quickest GPU-to-GPU communication. A single DGX B200 node with eight GPUs helps eight Presto GPU employees, all linked over NVLink 5.0 with 1,800 GB/s bidirectional bandwidth. GPU-accelerated Presto on DGX B200 reveals clear efficiency benefits over Presto CPU operating on 8-10 nodes of Intel Xeon 6642Y servers. 

This publish focuses on benchmarks derived from TPC-H, which embody 22 analytical queries, use parquet file knowledge sources, and change decimal sorts with float sorts for this demonstration. The measured question runtime contains SQL parsing, plan optimization, employee execution, and returning closing outcomes. Runtimes had been measured by operating every question 5 occasions and averaging the final 4 values. Benchmarking knowledge had been collected utilizing supply builds of Presto, Velox and cuDF. For complete construct, deployment, and benchmarking scripts, see the rapidsai/velox-testing GitHub repo.

Determine 1 compares the question runtime for multinode CPU Presto and single-node multi-GPU Presto. The Presto CPU configuration for scale issue 1K used Presto C++ employees, eight nodes of two-socket Intel Xeon 6642Y servers with 250 GiB system reminiscence, and a Lustre knowledge supply containing parquet information. At scale issue 3K, the Presto CPU configuration used 10 nodes of the identical servers. The Presto CPU configuration ran the benchmarks scorching, with the employees utilizing Velox async knowledge cache, and the coordinator utilizing comfortable affinity to extend cache utilization. The Presto GPU configuration used one NVIDIA DGX B200 with various variety of lively GPUs, and scorching cache parquet knowledge supply. 

At scale issue 1K (~1 TB knowledge set), GPU-accelerated Presto delivered 2.5x-8x decrease latency relying on the variety of lively GPUs. Presto GPU operating with one B200 GPU confirmed 2.5x sooner runtime in comparison with an eight-node Presto CPU cluster, and Presto GPU operating with eight B200 GPUs confirmed 8.2x sooner runtimes in comparison with an eight-node Presto CPU cluster. At scale issue 3K (~3 TB knowledge set), Presto GPU delivered 3x-8x decrease latency. Presto GPU operating with three B200 GPUs confirmed 3.6x sooner runtimes in comparison with a 10-node Presto CPU cluster, and Presto GPU operating with eight B200 GPUs confirmed 7.8x sooner runtime than a 10-node Presto CPU cluster. 

Bar chart showing Presto CPU performance with 8-10 nodes on scale factors 1K and 3K, and Presto GPU performance with 1 to 8 GPUs.
Bar chart showing Presto CPU performance with 8-10 nodes on scale factors 1K and 3K, and Presto GPU performance with 1 to 8 GPUs.
Determine 1. Question runtimes for TPC-H-derived scale elements 1K and 3K, evaluating Presto CPU multinode and Presto GPU multi-GPU single-node configurations 

How does GPU-accelerated Presto scale out on GB200 NVL72? 

GPU-accelerated Presto additionally scales out to a number of nodes with glorious noticed efficiency on NVLink-connected techniques comparable to NVIDIA GB200 NVL72. The NVIDIA GB200 NVL72 system features a whole of 18 nodes, the place every node contains two Grace CPUs, 4 B200 GPUs, and 4 ConnectX-7 (CX7) 400 Gbps community interface playing cards. The GPUs within the cluster are all linked by NVLink, opening the community for site visitors between compute and storage.

For GPU-accelerated Presto benchmarking, the NVL72 cluster was paired with IBM Storage Scale, an information storage system with 20 storage nodes, 10 PB capability and peak bandwidth of ~4.5 TiB/s. IBM Storage Scale, previously IBM Spectrum Scale or Normal Parallel File System (GPFS), is a clustered, POSIX-compliant parallel file system, offering native distant direct reminiscence entry (RDMA) transport over InfiniBand/RoCE. Collectively, NVIDIA GDS and IBM Storage Scale enable file knowledge to maneuver immediately from storage gadget to GPU reminiscence, bypassing host CPU and system-memory bounce buffers.

Determine 2 compares the full question runtime for TPC-H-derived scale elements 10K and 30K for the NVL72 GB200 cluster, highlighting essentially the most important I/O and communication optimizations throughout cluster carry up. The question runtime knowledge proven in Determine 2 used eight lively nodes out of 18 whole nodes, comparable to 32 Presto GPU employees collaborating within the analytical workloads. 

 Bar chart showing total runtime for “TPC-H derived” with 10K and 30K scale factor and node counts of eight.
 Bar chart showing total runtime for “TPC-H derived” with 10K and 30K scale factor and node counts of eight.
Determine 2. Scale out testing with Presto GPU from TPC-H derived scale issue 10K and 30K utilizing NVIDIA GB200 NVL72, displaying efficiency enhancements over a number of rounds of optimization

The primary-run situation used POSIX reads to IBM Storage Scale, unoptimized I/O parameters, and untuned UcxExchange configuration. The subsequent situation–gadget reads and 16 MiB I/O duties–delivered ~30% sooner runtime by enabling GDS and growing I/O job dimension from 4 MiB to 16 MiB, the beneficial dimension for IBM Storage Scale. 

The subsequent situation–plus 16 I/O threads–introduced an extra ~17% sooner runtime as a consequence of higher NVLink saturation when utilizing 16 I/O threads as an alternative of 4 I/O threads. The ultimate situation–plus rebatching and Q11 rewrite–lowered runtime one other 35% as a consequence of giant change batch sizes and fewer GPU idle time. Total the I/O and communication optimizations yielded 64% sooner question runtimes.

Q11 rewrite modified the SELECT assertion to an INSERT INTO assertion, which lowered runtime on Q11 from 50 seconds to 2 seconds. We noticed that Q11 as a SELECT assertion resulted in <5% GPU utilization time within the cluster, as a consequence of a bottleneck in sending question outcomes from a GPU employee to the coordinator over the default HttpExchange. Working Q11 as an INSERT INTO assertion used the GPU-based parquet author to effectively retailer ends in IBM Storage Scale. 

Shifting ahead, we plan to enhance the throughput of sending ends in Presto from employee to coordinator, conserving excessive GPU utilization with out adjusting queries. 

Tips on how to obtain sooner, light-weight I/O with GDS

On the NVL72 cluster with IBM Storage Scale, GDS is a important instrument for reaching peak efficiency and value efficiency. IBM Storage Scale allows two predominant knowledge paths: POSIX reads that stage knowledge in a consumer web page pool earlier than copying to GPU reminiscence, and GDS reads that populate GPU reminiscence utilizing RMDA. GDS is topology-aware, and ensures that the information path from the CX7 community card to GPU reminiscence stays throughout the identical NUMA node. The model of IBM Storage Scale on this research was not topology-aware for POSIX reads, so these reads incur each the additional staging buffer copy in addition to penalties from NUMA boundary crossing.

We analyzed TPC-H-derived 10K efficiency in two I/O configurations: POSIX chilly with a nominal web page pool dimension of fifty GiB, and GDS chilly, which bypasses caching. Working scale issue 10K with two nodes (eight GPUs), we noticed that GDS reads exhibit important benefits over POSIX reads. At two nodes and eight GPUs, GDS chilly reads confirmed ~2x sooner runtimes than POSIX chilly reads as a consequence of POSIX slowdowns from bounce buffer copying and NUMA boundary crossing. Because of the environment friendly knowledge path and lowered consumption of host sources, we count on GDS reads to be the popular strategy for efficiency and cost-performance on NVL72 with IBM Storage Scale.

Bar chart showing POSIX cold and GDS cold read for two nodes.
Bar chart showing POSIX cold and GDS cold read for two nodes.
Determine 3. POSIX chilly and GDS chilly reads for scale issue 10K on NVIDIA GB200 NVL72 for 2 nodes (eight GPUs)

Get began with GPU-accelerated Presto

Whether or not you’re operating interactive dashboards or nightly jobs, GPU-accelerated Presto supplies your analytical workloads with low latency and excessive throughput.

With integration within the IBM watsonx.knowledge platform, GPU-accelerated Presto is now prepared for testing in your manufacturing workloads. Register for entry to the technical preview of GPU-accelerated Presto on watsonx.knowledge. 

For builders and engineers inquisitive about Presto testing and deployment, check with the Presto Native gpu-nightly tag on the Presto DockerHub. To study extra, see Getting Began with GPU-Accelerated Presto C++.



Source link

Tags: AnalyticalGB200GPUAcceleratedLowLatencyNVIDIANVL72Prestorunningworkloads
Previous Post

Constructed to bounce again: How Azure resiliency developed

Next Post

The Decline of Deviance 2

Next Post
The Decline of Deviance 2

The Decline of Deviance 2

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Fetching latest news…
FUTURENEWS24
Live Feed
All
AI
Dev
Industry
Frontier
Updates in 60s
FN24 AI & Tech
View All →
Future News 24

The world's leading source for AI research, emerging technology, and the people building the future. Independent, rigorous, and always ahead.

CATEGORIES

  • AI Platforms & Apps
  • AI Research & Breakthroughs
  • BioTechnology
  • Data Science & MLOps
  • Decentralized Technology
  • Developer AI & Open-Source Ecosystem
  • Emerging Technologies & Innovations
  • Ethics & Policy
  • Industry & Business
  • Quantum Computing
  • Uncategorized

LATEST

  • [2602.13312] PeroMAS: A Multi-agent System of Perovskite Materials Discovery
  • GPT-6 Astra overview: code overview good points, privateness, and value
  • GPT-6 Astra: Options, Benchmarks, Pricing, and What’s New
  • About Us
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA 
  • Cookie Policy
  • Terms and Conditions
  • Contact us

© 2026 Future News 24. All rights reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized

© 2026 Future News 24. All rights reserved.

Website security powered by MilesWeb