{"id":1356,"date":"2026-06-22T16:00:00","date_gmt":"2026-06-22T16:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/"},"modified":"2026-06-23T01:59:30","modified_gmt":"2026-06-23T01:59:30","slug":"cccl-runtime-a-modern-c-runtime-for-cuda","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/","title":{"rendered":"CCCL Runtime: A Fashionable C++ Runtime for CUDA"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">The NVIDIA CUDA Core Compute Libraries (CCCL) gives pleasant and environment friendly abstractions for CUDA builders in C++ and Python. It options:<\/p>\n<p>Parallel algorithms \u2013 Host-launched algorithms together with type, scan and scale back that take away the necessity to write customized kernels for widespread operations\u00a0<\/p>\n<p>Cooperative algorithms \u2013 System-side algorithms corresponding to block-wide or warp-wide reductions or scans that simplify customized kernel improvement\u00a0<\/p>\n<p>Language idiomatic CUDA abstractions \u2013 Basic abstractions for CUDA-specific operations together with reminiscence allocation, useful resource administration, and {hardware} options\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">This publish introduces a brand new group of performance in CCCL that gives modernized C++ abstractions for elementary CUDA programming mannequin ideas that make CUDA C++ improvement safer and extra handy.\u00a0\u00a0<\/p>\n<h2 id=\"what_is_cccl_runtime\" class=\"wp-block-heading\">What&#8217;s CCCL runtime?<\/h2>\n<p class=\"wp-block-paragraph\">NVIDIA CCCL runtime is a brand new set of idiomatic C++ APIs that implement core CUDA performance: stream administration, reminiscence allocation, kernel launches, and extra.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The acquainted NVIDIA CUDA runtime was initially developed as a comfort layer on prime of the CUDA driver API. The brand new CCCL runtime goals to be another with the identical purpose, however with an up to date design aligned with fashionable C++. Determine 1, under, reveals the connection between the three CUDA API surfaces talked about above:<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a39e880a5282&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a39e880a5282\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"507\" height=\"144\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9.webp\" alt=\"Both CCCL runtime API and CUDA runtime API are built on top of the CUDA driver API. Both runtimes are also easily interoperable with each other&#10;\" class=\"wp-image-118838\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9.webp 507w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-179x51.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-300x85.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-500x142.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-160x45.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-362x103.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-387x110.png 387w\" sizes=\"(max-width: 507px) 100vw, 507px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"507\" height=\"144\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9.webp\" alt=\"Both CCCL runtime API and CUDA runtime API are built on top of the CUDA driver API. Both runtimes are also easily interoperable with each other&#10;\" class=\"lazyload wp-image-118838\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9.webp 507w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-179x51.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-300x85.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-500x142.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-160x45.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-362x103.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/image1-9-387x110.png 387w\" data-sizes=\"(max-width: 507px) 100vw, 507px\"\/><figcaption class=\"wp-element-caption\">Determine 1. Stack diagram of various CUDA API surfaces<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">CCCL runtime is a set of headers inside CCCL, corresponding to , , and . It leverages fashionable C++ options to supply extra handy and sturdy abstractions than what was attainable inside the C supply compatibility constraints of the standard CUDA runtime API.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">We additionally took the chance to include classes realized over 20 years of CUDA evolution into the API design. Even with all these modifications, CCCL runtime gives compatibility helpers that permit builders undertake it incrementally with out rewriting surrounding code that makes use of the CUDA runtime API.<\/p>\n<p class=\"wp-block-paragraph\">As CUDA applications develop extra complicated,\u00a0 with a number of libraries sharing units, streams, and reminiscence, the necessity for APIs that compose cleanly and make dependencies specific turns into extra urgent. That&#8217;s the house CCCL runtime is designed to fill.<\/p>\n<h2 id=\"the_code\" class=\"wp-block-heading\">The code<\/h2>\n<p class=\"wp-block-paragraph\">Right here is the basic vectorAdd instance carried out with the brand new CCCL runtime APIs. If you happen to\u2019ve written CUDA earlier than, the general construction shall be acquainted: Deal with what\u2019s totally different. Don\u2019t attempt to perceive all the things directly, the remainder of this publish will stroll by means of this instance to elucidate the semantics and design decisions behind CCCL runtime.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\n#embrace<br \/>\n#embrace<br \/>\n#embrace<br \/>\n#embrace<br \/>\n#embrace<br \/>\n#embrace <\/p>\n<p>struct kernel {<br \/>\n  template<br \/>\n  __device__ void operator()(Config config,<br \/>\n                             cuda::std::span A,<br \/>\n                             cuda::std::span B,<br \/>\n                             cuda::std::span C) {<br \/>\n    auto tid = cuda::gpu_thread.rank(cuda::grid, config);<br \/>\n    if (tid &lt; A.dimension())<br \/>\n      C[tid] = A[tid] + B[tid];<br \/>\n  }<br \/>\n};                                                                                                                                                                                                                                                                                  <\/p>\n<p>int foremost() {<br \/>\n  \/\/ 1. Gadgets and streams<br \/>\n  cuda::device_ref machine = cuda::units[0];<br \/>\n  cuda::stream stream{machine};                                     <\/p>\n<p>  \/\/ 2. Reminiscence allocation<br \/>\n  auto pool = cuda::device_default_memory_pool(machine);            <\/p>\n<p>  int num_elements = 1000;<br \/>\n  auto A = cuda::make_buffer(stream, pool, num_elements, 1);<br \/>\n  auto B = cuda::make_buffer(stream, pool, num_elements, 2);<br \/>\n  auto C = cuda::make_buffer(stream, pool, num_elements, cuda::no_init);                                                          <\/p>\n<p>  \/\/ 3. Kernel launch<br \/>\n  constexpr int threads_per_block = 256;<br \/>\n  auto config = cuda::distribute(num_elements); <\/p>\n<p>  cuda::launch(stream, config, kernel{}, A, B, C);                 <\/p>\n<p>  \/\/ Make the CPU thread look ahead to the GPU work to complete.<br \/>\n  stream.sync();<br \/>\n  return 0;<br \/>\n}\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">The instance may be damaged down into the next three foremost sections:<\/p>\n<h2 id=\"1_devices_and_streams\" class=\"wp-block-heading\">1.) Gadgets and streams<\/h2>\n<p class=\"wp-block-paragraph\">Contemplate the creation of a stream utilizing the CUDA Runtime API as the next code snippet reveals.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\ncudaStream_t stream;<br \/>\ncudaStreamCreate(&amp;stream); \/\/ related to whichever machine occurs to be &#8220;present&#8221;\n<\/div>\n<p class=\"wp-block-paragraph\">Notice this creates a stream, however the stream is related to whichever machine is present when cudaStreamCreate known as.\u00a0 Primarily based on this name alone, you don\u2019t know which machine the stream is related to.<\/p>\n<p class=\"wp-block-paragraph\">Distinction that with utilizing CCCL runtime API as illustrated by the code snippet that follows.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\ncuda::device_ref machine = cuda::units[0];<br \/>\ncuda::stream stream{machine};\n<\/div>\n<p class=\"wp-block-paragraph\">The above code snippet reveals the way to create a stream on a selected machine. The primary line illustrates a core design precept: CCCL runtime makes use of devoted sorts as a substitute of uncooked identifiers. A tool is a device_ref, not a plain integer; a stream is an object, not an opaque pointer. Robust typing throughout the API helps catch errors at compile time slightly than chasing them at runtime.\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The second line illustrates one other precept: making dependencies specific. In each CCCL runtime and the CUDA runtime API, a stream is related to a tool. The distinction is how. Right here, the cuda::stream constructor takes the machine as an specific argument whereas with the CUDA runtime API the stream is related to whichever machine is lively when the stream is created.<\/p>\n<p class=\"wp-block-paragraph\">Express dependencies allow native reasoning. You&#8217;ll be able to learn a perform and perceive what it does with out monitoring the worldwide state. Additionally they enhance composability: When a number of libraries are used, none of them want to avoid wasting and restore implicit state throughout calls to keep away from interfering with one another.\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">A associated consequence is that CCCL runtime doesn\u2019t expose the default stream. Managing the which means of the default stream requires monitoring the present machine, which is precisely the type of implicit state we&#8217;re shifting away from. Whereas a default stream from the CUDA runtime API can nonetheless be wrapped into CCCL runtime sorts, its utilization is discouraged; something involving the default stream must be dealt with by means of the CUDA runtime API instantly. With no default stream within the API, the notion of a \u201cblocking stream\u201d not applies, so all CCCL runtime streams are created as non-blocking.<\/p>\n<h3 id=\"resource_ownership_owning_types_and_refs\" class=\"wp-block-heading\">Useful resource possession: Proudly owning sorts and refs<\/h3>\n<p class=\"wp-block-paragraph\">Following the instance of std::string and std::string_view, many CUDA objects have two sorts in CCCL runtime: an proudly owning sort and a non-owning sort with a _ref suffix; cuda::stream owns the underlying cudaStream_t deal with and destroys it in its destructor. The cuda::stream_ref holds the deal with with out managing its lifetime and is trivially copyable.\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The _ref sorts are important for composability with current code. If a stream deal with\u2019s lifetime is managed elsewhere, cudaStream_t implicitly converts to cuda::stream_ref, and the uncooked deal with may be retrieved with .get(). To switch possession, cuda::stream::from_native_handle wraps a uncooked deal with into the proudly owning sort, and .launch() relinquishes possession again.\u00a0\u00a0<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nvoid stream_type_example(cudaStream_t deal with) {<br \/>\n  cuda::stream_ref non_owning{deal with};<br \/>\n  assert(deal with == non_owning.get());<\/p>\n<p>  cuda::stream proudly owning = cuda::stream::from_native_handle(deal with);<br \/>\n  assert(deal with == proudly owning.get());<br \/>\n  assert(deal with == proudly owning.launch());<br \/>\n}\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">The identical sample applies to occasions, reminiscence swimming pools, and different CUDA objects: cuda::device_ref has no proudly owning counterpart as a result of there is no such thing as a machine state to personal.<\/p>\n<h2 id=\"2_memory_allocation_\" class=\"wp-block-heading\">2.) Reminiscence allocation <\/h2>\n<div class=\"wp-block-syntaxhighlighter-code \">\nauto pool = cuda::device_default_memory_pool(machine);<\/p>\n<p>auto A = cuda::make_buffer(stream, pool, num_elements, 1);<br \/>\nauto B = cuda::make_buffer(stream, pool, num_elements, 2);<br \/>\nauto C = cuda::make_buffer(stream, pool, num_elements, cuda::no_init);\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">The following part demonstrates asynchronously allocating and initializing machine reminiscence. Right here we see the following design precept: APIs are asynchronous by default. Slightly than distinguishing synchronous and asynchronous variants by title, CCCL runtime makes use of a easy conference: If an API takes a stream as its first argument, it operates in stream order. We don\u2019t plan to supply synchronous counterparts for APIs which have each variants within the CUDA runtime API. \u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Reminiscence allocation is the place this issues most in apply. Stream-ordered reminiscence administration through reminiscence swimming pools has been obtainable since CUDA 11.2 (defined right here), and CUDA 13.0 expanded it to managed and host reminiscence. Reminiscence pooling and fewer frequent synchronization factors are generally important to succeed in most efficiency, and stream-ordered reminiscence administration composes naturally with the remainder of the asynchronous programming mannequin. To convey these tips, CCCL runtime makes reminiscence swimming pools and stream-ordered allocation the default. On older CUDA variations and platforms, the place newer reminiscence pool sorts usually are not but supported, we offer non-stream-ordered allocation as a fallback, however plan to take away it as soon as pool assist is common.\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Within the snippet above, we first question the default reminiscence pool for a given machine, passing it as an specific argument slightly than counting on cudaMallocAsync\u2018s implicit machine choice. The instance makes use of the default pool which must be most well-liked the place attainable, however CCCL runtime additionally permits creating separate pool objects when totally different pool settings are wanted.<\/p>\n<p class=\"wp-block-paragraph\">The pool reference is then used to create three buffers utilizing the brand new cuda::make_buffer. It takes a stream as its first argument to sign stream-ordered operation. Every buffer submits three operations to that stream: allocation from the required pool, initialization, and ultimately deallocation when the buffer goes out of scope.<\/p>\n<p class=\"wp-block-paragraph\">Initialization is necessary until explicitly opted out with cuda::no_init, as with buffer C which shall be overwritten by the kernel. Uninitialized machine reminiscence is a typical supply of hard-to-diagnose bugs, so we selected to require an specific opt-out slightly than making it the silent default. Enter buffers A and B have all components initialized to 1 and a couple of, respectively. Buffers assist further initialization modes as effectively, for instance from one other buffer or a spread.<\/p>\n<h3 id=\"buffer_lifetime_and_deallocation\" class=\"wp-block-heading\">Buffer lifetime and deallocation<\/h3>\n<p class=\"wp-block-paragraph\">The stream handed to make_buffer is saved contained in the buffer and used for deallocation when the buffer is destroyed. This implies the buffer ought to typically maintain the stream that corresponds to its utilization, in order that computation is correctly ordered with deallocation. It&#8217;s attainable to alter the stream later with .set_stream() or manually set off destruction on a selected stream with .destroy(), however the default habits is designed to do the appropriate factor within the widespread case.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\n{<br \/>\n  auto pool = cuda::device_default_memory_pool(machine);<br \/>\n  \/\/ Equal to cudaMallocFromPoolAsync on the stream, probably together with initialization pushed into the stream as effectively. Saves the stream for future deallocation<br \/>\n  auto buffer = cuda::make_buffer(allocation_stream, pool, \/*&#8230; *\/);<\/p>\n<p>  \/\/ buffer utilization&#8230;<br \/>\n}<br \/>\n\/\/ Closing bracket will name cudaFreeAsync on allocation_stream, there may be additionally buffer.destroy(which_stream) to maintain the habits specific\n<\/p><\/div>\n<h2 id=\"3_kernel_launch\" class=\"wp-block-heading\">3.) Kernel launch<\/h2>\n<div class=\"wp-block-syntaxhighlighter-code \">\nstruct kernel {<br \/>\n  template<br \/>\n  __device__ void operator()(Config config,<br \/>\n                             cuda::std::span A,<br \/>\n                             cuda::std::span B,<br \/>\n                             cuda::std::span C) {<br \/>\n    auto tid = cuda::gpu_thread.rank(cuda::grid, config);<br \/>\n    if (tid &lt; A.dimension())<br \/>\n      C[tid] = A[tid] + B[tid];<br \/>\n  }<br \/>\n};<\/p>\n<p>\/\/ &#8230;<\/p>\n<p>constexpr int threads_per_block = 256;<br \/>\nauto config = cuda::distribute(num_elements);<\/p>\n<p>cuda::launch(stream, config, kernel{}, A, B, C);\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">The ultimate part demonstrates configuring and launching the kernel on the GPU with cuda::launch.\u00a0\u00a0\u00a0\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">cuda::launch takes three teams of arguments:<\/p>\n<p>The stream to run on<\/p>\n<p>A configuration object that encodes the thread hierarchy (block and grid sizes) together with different launch choices. Right here, cuda::distribute creates a configuration that launches at the least num_elements threads grouped into blocks of threads_per_block. This replaces the widespread sample many CUDA builders are accustomed to of (N + block_size &#8211; 1) \/ block_size<\/p>\n<p>The\u00a0 kernel and its arguments<\/p>\n<h2 id=\"compile-time_configuration_flow\" class=\"wp-block-heading\">Compile-time configuration stream<\/h2>\n<p class=\"wp-block-paragraph\">Probably the most novel facet of cuda::launch is the way it strikes compile-time data from the host launch web site into machine code by means of the kind system. For instance, discover how the\u00a0block dimension is offered as a template argument to cuda::distribute, which suggests it&#8217;s encoded within the configuration object\u2019s sort. <\/p>\n<p class=\"wp-block-paragraph\">When the kernel accepts that configuration as its first argument, cuda::launch passes it by means of mechanically. Contained in the kernel, this static data is accessible once we compute the rank of the calling thread contained in the grid:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nauto tid = cuda::gpu_thread.rank(cuda::grid, config);\n<\/div>\n<p class=\"wp-block-paragraph\">As a result of the block dimension is thought at compile time, the rank calculation can use solely the x dimension and skip the runtime block-size question solely. It is a easy instance, however the mechanism generalizes.\u00a0 The CCCL documentation reveals additional circumstances the place configuration-embedded data is used to specialize machine code.<\/p>\n<p>Typically kernel implementation makes assumptions concerning the precise form of the grid and\/or block. Compile time data within the configuration object permits kernel authors to implement checks to make sure alignment of the kernel and the decision web site in these circumstances.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\ntemplate<br \/>\n__global__ void kernel(Config conf) {<br \/>\n    \/\/ Be sure that the block is one dimensional with 256 threads<br \/>\n    static_assert(cuda::gpu_thread.static_dims(cuda::block, conf).x == 256);<br \/>\n    static_assert(cuda::gpu_thread.static_dims(cuda::block, conf).y == 1);<br \/>\n    static_assert(cuda::gpu_thread.static_dims(cuda::block, conf).z == 1);<br \/>\n}\n<\/div>\n<h2 id=\"kernel_functors\" class=\"wp-block-heading\">Kernel functors<\/h2>\n<p class=\"wp-block-paragraph\">You might have observed the kernel is a struct with a __device__ operator() slightly than a __global__ perform. Whereas cuda::launch helps current __global__ capabilities, we additionally launched kernel functors: sorts with a __device__-annotated name operator. The sensible benefit is that template arguments are deduced mechanically, whereas __global__ capabilities used with cuda::launch require specific instantiation.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\ntemplate<br \/>\n__global__ void kernel_function(T enter) {<br \/>\n  \/\/ physique &#8230;<br \/>\n}<\/p>\n<p>struct kernel_functor {<br \/>\n  template<br \/>\n  __device__ void operator()(T enter) {<br \/>\n  \/\/ physique &#8230;<br \/>\n  }<br \/>\n};<\/p>\n<p>\/\/ specific template instantiation is required with a __global__ perform<br \/>\ncuda::launch(stream, config, kernel_function, 42);<br \/>\n\/\/ deduction from arguments for a functor with __device__ name operator<br \/>\ncuda::launch(stream, config, kernel_functor{}, 42);\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">That is what makes the compile-time configuration stream work. The config template parameter is deduced from the configuration object handed by cuda::launch. Kernel functors additionally cowl machine lambdas and have further options described within the CCCL documentation.<\/p>\n<h2 id=\"automatic_argument_transformation\" class=\"wp-block-heading\">Computerized argument transformation<\/h2>\n<p class=\"wp-block-paragraph\">cuda::buffer owns its underlying allocation, however CUDA kernels can solely settle for trivially copyable arguments. When a buffer is handed to cuda::launch, it&#8217;s mechanically reworked to cuda::std::span. There isn&#8217;t any must manually assemble the span or extract a uncooked pointer. The kernel signature displays how the info is definitely used on the machine aspect.<\/p>\n<h2 id=\"what\u2019s_next\" class=\"wp-block-heading\">What\u2019s subsequent<\/h2>\n<p class=\"wp-block-paragraph\">This publish lined the core concepts behind CCCL runtime: specific dependencies, robust typing, asynchronous-by-default APIs, and clear interoperability with current CUDA code. However a walkthrough of 1 instance can solely present a lot. The CCCL documentation has extra detailed protection of every API, together with further buffer initialization modes, occasion administration, knowledge motion, and superior kernel launch options like dynamic shared reminiscence and different launch attributes. CCCL runtime is accessible as we speak in CCCL. We\u2019d love to listen to your suggestions as you attempt it out.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/cccl-runtime-a-modern-c-runtime-for-cuda\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The NVIDIA CUDA Core Compute Libraries (CCCL) gives pleasant and environment friendly abstractions for CUDA builders in C++ and Python. It options: Parallel algorithms \u2013 Host-launched algorithms together with type, scan and scale back that take away the necessity to write customized kernels for widespread operations\u00a0 Cooperative algorithms \u2013 System-side algorithms corresponding to block-wide or [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1358,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/featured-image-JB-616.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[1793,1795,1083,1794],"class_list":["post-1356","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-cccl","tag-cuda","tag-modern","tag-runtime"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>CCCL Runtime: A Fashionable C++ Runtime for CUDA - Future News 24<\/title>\n<meta name=\"description\" content=\"The NVIDIA CUDA Core Compute Libraries (CCCL) provides delightful and efficient abstractions for CUDA developers in C++ and Python. It features: This post&#8230;\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"CCCL Runtime: A Fashionable C++ Runtime for CUDA - Future News 24\" \/>\n<meta property=\"og:description\" content=\"The NVIDIA CUDA Core Compute Libraries (CCCL) provides delightful and efficient abstractions for CUDA developers in C++ and Python. It features: This post&#8230;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-22T16:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-23T01:59:30+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/featured-image-JB-616.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/featured-image-JB-616.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"12 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"CCCL Runtime: A Fashionable C++ Runtime for CUDA\",\"datePublished\":\"2026-06-22T16:00:00+00:00\",\"dateModified\":\"2026-06-23T01:59:30+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/\"},\"wordCount\":2439,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/featured-image-JB-616.webp\",\"keywords\":[\"CCCL\",\"CUDA\",\"modern\",\"Runtime\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/\",\"name\":\"CCCL Runtime: A Fashionable C++ Runtime for CUDA - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/featured-image-JB-616.webp\",\"datePublished\":\"2026-06-22T16:00:00+00:00\",\"dateModified\":\"2026-06-23T01:59:30+00:00\",\"description\":\"The NVIDIA CUDA Core Compute Libraries (CCCL) provides delightful and efficient abstractions for CUDA developers in C++ and Python. It features: This post&#8230;\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/featured-image-JB-616.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/featured-image-JB-616.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/22\\\/cccl-runtime-a-modern-c-runtime-for-cuda\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"CCCL Runtime: A Fashionable C++ Runtime for CUDA\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"CCCL Runtime: A Fashionable C++ Runtime for CUDA - Future News 24","description":"The NVIDIA CUDA Core Compute Libraries (CCCL) provides delightful and efficient abstractions for CUDA developers in C++ and Python. It features: This post&#8230;","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/","og_locale":"en_US","og_type":"article","og_title":"CCCL Runtime: A Fashionable C++ Runtime for CUDA - Future News 24","og_description":"The NVIDIA CUDA Core Compute Libraries (CCCL) provides delightful and efficient abstractions for CUDA developers in C++ and Python. It features: This post&#8230;","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/","og_site_name":"Future News 24","article_published_time":"2026-06-22T16:00:00+00:00","article_modified_time":"2026-06-23T01:59:30+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/featured-image-JB-616.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/featured-image-JB-616.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"12 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"CCCL Runtime: A Fashionable C++ Runtime for CUDA","datePublished":"2026-06-22T16:00:00+00:00","dateModified":"2026-06-23T01:59:30+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/"},"wordCount":2439,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/featured-image-JB-616.webp","keywords":["CCCL","CUDA","modern","Runtime"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/","name":"CCCL Runtime: A Fashionable C++ Runtime for CUDA - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/featured-image-JB-616.webp","datePublished":"2026-06-22T16:00:00+00:00","dateModified":"2026-06-23T01:59:30+00:00","description":"The NVIDIA CUDA Core Compute Libraries (CCCL) provides delightful and efficient abstractions for CUDA developers in C++ and Python. It features: This post&#8230;","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/featured-image-JB-616.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/06\/featured-image-JB-616.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/22\/cccl-runtime-a-modern-c-runtime-for-cuda\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"CCCL Runtime: A Fashionable C++ Runtime for CUDA"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1356","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=1356"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1356\/revisions"}],"predecessor-version":[{"id":1357,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1356\/revisions\/1357"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/1358"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=1356"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=1356"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=1356"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}