{"id":2732,"date":"2026-07-22T16:35:00","date_gmt":"2026-07-22T16:35:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/"},"modified":"2026-07-23T09:59:16","modified_gmt":"2026-07-23T09:59:16","slug":"make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/","title":{"rendered":"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">A TensorRT engine construct can take seconds to many minutes. Giant strongly typed fashions, deep tactic search, and a chilly timing cache on a brand-new GPU SKU can depart builders, finish customers, or AI brokers looking at a frozen terminal with no concept whether or not to attend, retry, or kill the method. Most NVIDIA TensorRT integrations report nothing throughout a construct or present no technique to abort early. In a long-running agent workflow, this turns into wasted GPU-hours and caught classes.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a61e5f3dc003&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a61e5f3dc003\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1536\" height=\"1024\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65.webp\" alt=\"Dark terminal showing four nested ASCII progress bars in NVIDIA green, labeled &quot;Building Engine,&quot; &quot;Tactic Selection,&quot; &quot;Timing Cache Warmup,&quot; and &quot;Kernel Autotune.&quot; Each row is indented to indicate nesting under its parent phase.\" class=\"wp-image-120286\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-173x115.png 173w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-300x200.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-768x512.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-625x417.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-645x430.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-450x300.png 450w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-135x90.png 135w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-362x241.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-165x110.png 165w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-1024x683.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-810x540.png 810w\" sizes=\"(max-width: 1536px) 100vw, 1536px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1536\" height=\"1024\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65.webp\" alt=\"Dark terminal showing four nested ASCII progress bars in NVIDIA green, labeled &quot;Building Engine,&quot; &quot;Tactic Selection,&quot; &quot;Timing Cache Warmup,&quot; and &quot;Kernel Autotune.&quot; Each row is indented to indicate nesting under its parent phase.\" class=\"lazyload wp-image-120286\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-173x115.png 173w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-300x200.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-768x512.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-625x417.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-645x430.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-450x300.png 450w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-135x90.png 135w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-362x241.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-165x110.png 165w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-1024x683.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-65-810x540.png 810w\" data-sizes=\"(max-width: 1536px) 100vw, 1536px\"\/><figcaption class=\"wp-element-caption\">Determine 1. Stay, nested progress bars for a TensorRT engine construct, rendered by an IProgressMonitor subclass<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">TensorRT offers IProgressMonitor, an API for fixing this difficulty, and it has been in NvInfer.h for a number of releases. This tutorial walks via a minimal drop-in implementation for Python and C++, provides a cancel path that responds to Ctrl-C or a programmatic cease sign from an outer occasion loop, and reveals the place to floor the ensuing progress stream so an IDE, a service, or an agent runtime can use it.<\/p>\n<p class=\"wp-block-paragraph\">Each code block on this submit is lifted from or modeled on two NVIDIA-maintained OSS samples:<\/p>\n<p>Python: samples\/python\/simple_progress_monitor\/ (ResNet-50, strongly typed community)<\/p>\n<p>C++: samples\/sampleProgressMonitor\/ (MNIST)<\/p>\n<h2 id=\"what_iprogressmonitor_gives_you\" class=\"wp-block-heading\">What IProgressMonitor offers you<\/h2>\n<p class=\"wp-block-paragraph\">IProgressMonitor is an summary base class that TensorRT calls throughout the engine construct. You subclass it and override three strategies. The form is an identical in Python and C++; solely the spelling differs.<\/p>\n<figure class=\"wp-block-table\">ConceptPython methodC++ methodWhat you doPhase enteredphase_start(phase_name, parent_phase, num_steps)phaseStart(phaseName, parentPhase, nbSteps)Reserve a progress row and file num_steps.Step inside part completestep_complete(phase_name, step) -&gt; boolstepComplete(phaseName, step) -&gt; boolAdvance the bar. Return False\/false to cancel the construct.Part exitedphase_finish(phase_name)phaseFinish(phaseName)Tear down the row.<figcaption class=\"wp-element-caption\">Desk 1. The IProgressMonitor interface mirrored throughout Python and C++. The three strategies have an identical semantics, and step_complete is the one callback whose return worth adjustments the builder\u2019s conduct<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">A part whose parent_phase is non-null is nested inside one other part, so the monitor sees a tree of progress moderately than a flat record. The implementation have to be thread-safe as a result of TensorRT can name the identical monitor occasion from a number of inside threads.<\/p>\n<p class=\"wp-block-paragraph\">Wire the monitor to the builder by setting it on the IBuilderConfig. It&#8217;s a single name in both language:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nconfig.progress_monitor = MyMonitor()      # Python<\/p>\n<p>config-&gt;setProgressMonitor(&amp;myMonitor);     \/\/ C++\n<\/p><\/div>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a61e5f3dd047&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a61e5f3dd047\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1536\" height=\"1024\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67.webp\" alt=\"UML sequence diagram between the TensorRT Builder and a user IProgressMonitor subclass showing nested phase_start calls, step_complete returning true (continue) or false (cancel), and a red-highlighted cancel path where the builder unwinds by calling phase_finish early on every active phase.\" class=\"wp-image-120288\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-173x115.png 173w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-300x200.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-768x512.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-625x417.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-645x430.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-450x300.png 450w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-135x90.png 135w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-362x241.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-165x110.png 165w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-1024x683.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-810x540.png 810w\" sizes=\"(max-width: 1536px) 100vw, 1536px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1536\" height=\"1024\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67.webp\" alt=\"UML sequence diagram between the TensorRT Builder and a user IProgressMonitor subclass showing nested phase_start calls, step_complete returning true (continue) or false (cancel), and a red-highlighted cancel path where the builder unwinds by calling phase_finish early on every active phase.\" class=\"lazyload wp-image-120288\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-173x115.png 173w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-300x200.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-768x512.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-625x417.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-645x430.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-450x300.png 450w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-135x90.png 135w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-362x241.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-165x110.png 165w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-1024x683.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-67-810x540.png 810w\" data-sizes=\"(max-width: 1536px) 100vw, 1536px\"\/><figcaption class=\"wp-element-caption\">Determine 2. The callback sequence TensorRT drives throughout a construct, with the cancel path highlighted in purple<\/figcaption><\/figure>\n<\/div>\n<p class=\"wp-block-paragraph\">Learn the diagram from prime to backside. The builder opens the Constructing Engine part with phase_start, then opens Tactic Choice nested inside it with its parent_phase pointing again at Constructing Engine. Because the construct proceeds, the builder calls step_complete (the stable arrows) and your monitor returns a Boolean (the dashed arrows): true lets the construct proceed and false requests cancellation. Within the run proven right here, the monitor returns false at step 47, which is the purple cancel path, and the builder stops issuing new steps and unwinds. It calls phase_finish early on Tactic Choice after which on Constructing Engine, closing each lively part in reverse order.<\/p>\n<h2 id=\"what_this_tutorial_builds\" class=\"wp-block-heading\">What this tutorial builds<\/h2>\n<p class=\"wp-block-paragraph\">This tutorial reveals find out how to implement IProgressMonitor in Python and C++, add cancellation via step_complete, and route progress updates to a terminal, IDE, service, or agent runtime.<\/p>\n<h2 id=\"prerequisites\" class=\"wp-block-heading\">Conditions<\/h2>\n<p>One NVIDIA GPU.<\/p>\n<p>TensorRT (present OSS launch) and its Python bindings, or a construct of the C++ samples.<\/p>\n<p>Python 3.10 or newer (Python path).<\/p>\n<p>The TensorRT pattern information: ResNet-50 ONNX for Python and MNIST ONNX for C++. Each ship with the sample-data archive or are mounted beneath \/usr\/src\/tensorrt\/information within the official NGC containers.<\/p>\n<p>A terminal that helps ANSI virtual-terminal escapes. Any trendy Linux shell qualifies; Home windows Terminal works if VT is enabled.<\/p>\n<h3 id=\"1_subclass_iprogressmonitor_in_python\" class=\"wp-block-heading\">1. Subclass IProgressMonitor in Python<\/h3>\n<p class=\"wp-block-paragraph\">The subclass is small. It solely tracks which phases are lively and what number of steps every part incorporates.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nimport tensorrt as trt<br \/>\nfrom dataclasses import dataclass, discipline<br \/>\nfrom threading import Lock<\/p>\n<p>@dataclass<br \/>\nclass _PhaseState:<br \/>\n    num_steps: int<br \/>\n    current_step: int = 0<br \/>\n    guardian: str | None = None<\/p>\n<p>class RichProgressMonitor(trt.IProgressMonitor):<br \/>\n    def __init__(self):<br \/>\n        tremendous().__init__()<br \/>\n        self._lock = Lock()<br \/>\n        self._phases: dict[str, _PhaseState] = {}<br \/>\n        self._cancelled = False<br \/>\n\t     self._rendered_lines = 0<\/p>\n<p>    def phase_start(self, phase_name, parent_phase, num_steps):<br \/>\n        with self._lock:<br \/>\n        \tself._phases[phase_name] = _PhaseState(<br \/>\n            \tnum_steps=num_steps, guardian=parent_phase<br \/>\n        \t)<br \/>\n        \tself._render()<\/p>\n<p>    def step_complete(self, phase_name, step) -&gt; bool:<br \/>\n        with self._lock:<br \/>\n        \tif phase_name in self._phases:<br \/>\n                self._phases[phase_name].current_step = step<br \/>\n        \tself._render()<br \/>\n        \treturn not self._cancelled<\/p>\n<p>    def phase_finish(self, phase_name):<br \/>\n        with self._lock:<br \/>\n        \tself._phases.pop(phase_name, None)<br \/>\n        \tself._render()\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">Two issues to note. First, the Lock isn&#8217;t optionally available. TensorRT will name into the monitor from a number of inside threads, and rendering from a thread that doesn\u2019t personal the state will tear the show. Second, step_complete is the one callback that may cease the construct. phase_start returns None, so you can not reject a part earlier than it begins. The earliest cancellation level is the primary step_complete of that part.<\/p>\n<h3 id=\"2_render_nested_progress_bars_with_virtual-terminal_escapes\" class=\"wp-block-heading\">2. Render nested progress bars with virtual-terminal escapes<\/h3>\n<p class=\"wp-block-paragraph\">The renderer is the half that varies most by surroundings, so this part offers the form and factors to the upstream pattern for the production-grade implementation. The sample is:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\ndef _render(self):<br \/>\n    # Order phases by nesting depth so youngsters draw beneath dad and mom.<br \/>\n    rows = sorted(<br \/>\n        self._phases.objects(),<br \/>\n        key=lambda kv: (kv[1].guardian or &#8220;&#8221;, kv[0]),<br \/>\n    )<br \/>\n    # Transfer the cursor up by the variety of traces the PREVIOUS render printed,<br \/>\n    # not the present row depend \u2014 phases are added on nesting and eliminated on<br \/>\n    # phase_finish, so the 2 differ precisely when the tree adjustments form.<br \/>\n    if self._rendered_lines:<br \/>\n        print(f&#8221;x1b[{self._rendered_lines}A&#8221;, end=&#8221;&#8221;)<br \/>\n    for name, st in rows:<br \/>\n        # step is a 0-based index in [0, num_steps); +1 turns it into a<br \/>\n        # completed count so the bar can actually reach 100%.<br \/>\n        done = min(st.current_step + 1, st.num_steps)<br \/>\n        pct = done \/ max(st.num_steps, 1)<br \/>\n        bar = &#8220;\u2588&#8221; * int(40 * pct) + &#8220;\u00b7&#8221; * (40 &#8211; int(40 * pct))<br \/>\n        indent = &#8221;  &#8221; if st.parent else &#8220;&#8221;<br \/>\n        print(f&#8221;x1b[2K{indent}{name:&lt;28} [{bar}] {performed}\/{st.num_steps}&#8221;)<br \/>\n    # Clear rows left behind when a part finishes and the depend shrinks.<br \/>\n    for _ in vary(self._rendered_lines &#8211; len(rows)):<br \/>\n        print(&#8220;x1b[2K&#8221;)<br \/>\n    self._rendered_lines = len(rows)\n<\/div>\n<p class=\"wp-block-paragraph\">The upstream simple_progress_monitor.py renders the same shape with improved color and width handling. The escape sequence x1b[NA moves the cursor up N lines, and x1b[2K clears a line. The first render call writes blank rows; subsequent calls overwrite them in place.<\/p>\n<p class=\"wp-block-paragraph\">When this monitor is attached, do not redirect stdout to a file or pipe. The escape codes will be written verbatim into the log and make it unreadable. For non-terminal sinks, replace _render() with a structured emitter.<\/p>\n<h3 id=\"3_add_a_cancel_path\" class=\"wp-block-heading\">3. Add a cancel path<\/h3>\n<p class=\"wp-block-paragraph\">Cancellation is a three-line addition once the monitor exists. Install a SIGINT handler that flips the flag, then let step_complete honor it.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nimport signal<\/p>\n<p>def install_cancel(monitor: RichProgressMonitor):<br \/>\n    def handler(signum, frame):<br \/>\n        monitor._cancelled = True<br \/>\n        print(&#8220;nCancelling TensorRT build at next step boundary&#8230;&#8221;)<\/p>\n<p>    signal.signal(signal.SIGINT, handler)\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">Wire the monitor and run the builder:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nbuilder = trt.Builder(TRT_LOGGER)<br \/>\nnetwork = builder.create_network(<br \/>\n    1 &lt;&lt; int(trt.NetworkDefinitionCreationFlag.STRONGLY_TYPED)<br \/>\n)<\/p>\n<p>parser = trt.OnnxParser(network, TRT_LOGGER)<\/p>\n<p>with open(onnx_path, &#8220;rb&#8221;) as f:<br \/>\n    parser.parse(f.read())<\/p>\n<p>config = builder.create_builder_config()<\/p>\n<p>monitor = RichProgressMonitor()<br \/>\nconfig.progress_monitor = monitor<\/p>\n<p>install_cancel(monitor)<\/p>\n<p>serialized = builder.build_serialized_network(network, config)<\/p>\n<p>if serialized is None:<br \/>\n    if monitor._cancelled:<br \/>\n        print(&#8220;Build cancelled cleanly.&#8221;)<br \/>\n    else:<br \/>\n        print(&#8220;Build failed.&#8221;)\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">build_serialized_network() returns None on cancellation. The builder unwinds at the next step boundary, usually quickly, but not instantaneously, especially inside a long tactic-search step.<\/p>\n<p class=\"wp-block-paragraph\">Applications should surface cancellation latency to users. A simple \u201cCancelling\u2026\u201d message during the unwind window goes a long way.<\/p>\n<p class=\"wp-block-paragraph\">The same flag can be set from any non-signal path, such as an IDE Stop button, an agent timeout, or a CI cancel webhook. Set monitor._cancelled = True, and the build aborts at the next step boundary.<\/p>\n<h3 id=\"4_the_same_pattern_in_c++\" class=\"wp-block-heading\">4. The same pattern in C++<\/h3>\n<div class=\"wp-block-syntaxhighlighter-code \">\n#include<br \/>\n#include<br \/>\n#include<br \/>\n#include <\/p>\n<p>class RichProgressMonitor : public nvinfer1::IProgressMonitor {<br \/>\npublic:<br \/>\n    void phaseStart(char const* phaseName,<br \/>\n                    char const* parentPhase,<br \/>\n                    int32_t nbSteps) noexcept override {<br \/>\n        std::lock_guard g(mu_);<br \/>\n        phases_[phaseName] = {nbSteps, 0, parentPhase ? parentPhase : &#8220;&#8221;};<br \/>\n        render();<br \/>\n    }<\/p>\n<p>    bool stepComplete(char const* phaseName,<br \/>\n                      int32_t step) noexcept override {<br \/>\n        std::lock_guard g(mu_);<br \/>\n        auto it = phases_.discover(phaseName);<br \/>\n        if (it != phases_.finish())<br \/>\n            it-&gt;second.present = step;<br \/>\n        render();<br \/>\n        return !cancelled_.load();<br \/>\n    }<\/p>\n<p>    void phaseFinish(char const* phaseName) noexcept override {<br \/>\n        std::lock_guard g(mu_);<br \/>\n        phases_.erase(phaseName);<br \/>\n        render();<br \/>\n    }<\/p>\n<p>    void requestCancel() noexcept {<br \/>\n        cancelled_.retailer(true);<br \/>\n    }<\/p>\n<p>non-public:<br \/>\n    struct Part {<br \/>\n        int32_t nbSteps;<br \/>\n        int32_t present;<br \/>\n        std::string guardian;<br \/>\n    };<\/p>\n<p>    std::mutex mu_;<br \/>\n    std::unordered_map phases_;<br \/>\n    std::atomic cancelled_{false};<\/p>\n<p>    void render() noexcept;<br \/>\n};\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">Connect it the identical method:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\nauto config =<br \/>\n    std::unique_ptr(<br \/>\n        builder-&gt;createBuilderConfig());<\/p>\n<p>RichProgressMonitor monitor;<\/p>\n<p>config-&gt;setProgressMonitor(&amp;monitor);\n<\/p><\/div>\n<p class=\"wp-block-paragraph\">std::atomic for the cancel flag issues as a result of requestCancel() could also be referred to as from one other thread or a sign handler. Every little thing else mirrors the Python model.<\/p>\n<h2 id=\"where_to_wire_it_in_real_systems\" class=\"wp-block-heading\">The place to wire it in actual methods<\/h2>\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a61e5f3de99f&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a61e5f3de99f\" class=\"wp-block-image size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1536\" height=\"1024\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68.webp\" alt=\"Three-layer architecture diagram with the TensorRT Builder at the top, IProgressMonitor in the middle, and four application sinks below: terminal renderer, IDE extension, FastAPI service, and agent runtime. A dashed red arrow shows the cancel signal flowing from a sink back through the monitor to the builder.\" class=\"wp-image-120289\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-173x115.png 173w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-300x200.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-768x512.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-625x417.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-645x430.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-450x300.png 450w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-135x90.png 135w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-362x241.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-165x110.png 165w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-1024x683.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-810x540.png 810w\" sizes=\"(max-width: 1536px) 100vw, 1536px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1536\" height=\"1024\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68.webp\" alt=\"Three-layer architecture diagram with the TensorRT Builder at the top, IProgressMonitor in the middle, and four application sinks below: terminal renderer, IDE extension, FastAPI service, and agent runtime. A dashed red arrow shows the cancel signal flowing from a sink back through the monitor to the builder.\" class=\"lazyload wp-image-120289\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-173x115.png 173w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-300x200.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-768x512.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-625x417.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-645x430.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-450x300.png 450w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-135x90.png 135w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-362x241.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-165x110.png 165w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-1024x683.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-68-810x540.png 810w\" data-sizes=\"(max-width: 1536px) 100vw, 1536px\"\/><\/p>\n<\/figure>\n<p class=\"wp-block-paragraph\">Determine 3. IProgressMonitor is the only integration level between the builder and an software\u2019s surfaces<\/p>\n<p class=\"wp-block-paragraph\">The cancel arrow is drawn from the agent runtime for concreteness, however the identical mechanism applies to each sink. A Ctrl-C from the terminal, an IDE Cease button, an HTTP cancel webhook, or an agent timeout all flip the identical monitor._cancelled flag, and the cancel takes impact on the subsequent step_complete return.The place to wire it in actual methods<\/p>\n<p class=\"wp-block-paragraph\">The terminal is the straightforward case. The fascinating integrations route progress someplace else:<\/p>\n<p>IDE extension: Override _render() to emit $\/progress notifications within the Language Server Protocol, or equal window\/showProgress in protocol. Every part turns into one progress token; step_complete() turns into a report message; phase_finish() turns into finish.<\/p>\n<p>FastAPI \/ HTTP service: Run the construct on a background thread, and have _render() push entries into an asyncio.Queue that the request handler drains by way of Server-Despatched Occasions. The shopper will get a dwell stream; the cancel hook is only a POST \/builds\/{id}\/cancel that calls monitor.requestCancel().<\/p>\n<p>Agent device name: Emit one structured chunk per part transition ({&#8220;part&#8221;: &#8230;, &#8220;step&#8221;: &#8230;, &#8220;complete&#8221;: &#8230;}) into the tool-call stream. The agent runtime renders it within the user-visible hint, and the identical requestCancel() hook is what an agent timeout calls when the construct exceeds the finances. This sample additionally issues for agent runtimes. Lengthy-running builds must be observable and cancelable so brokers can report progress, implement time budgets, and cease cleanly.<\/p>\n<p class=\"wp-block-paragraph\">In all three circumstances, IProgressMonitor is the fitting boundary. Something above it (rendering, streaming, transport) is application-level; something beneath it (tactic timing, kernel choice) is the builder\u2019s enterprise.<\/p>\n<h2 id=\"edge_cases_to_handle\" class=\"wp-block-heading\">Edge circumstances to deal with<\/h2>\n<p class=\"wp-block-paragraph\">These behaviors are frequent sources of integration bugs:<\/p>\n<p>Don&#8217;t redirect stdout whereas the terminal renderer is hooked up. The escape sequences will pollute the log. For non-interactive sinks, swap the renderer for a structured emitter.<\/p>\n<p>phase_start() can\u2019t cancel. It returns None. The earliest cancel level is the primary step_complete() of that part. If the consumer cancels throughout an extended phase_start(), the construct will proceed till step one boundary.<\/p>\n<p>phase_finish() might hearth earlier than all num_steps are reported. This may occur throughout error restoration, builder-internal short-circuits, or when step_complete() returns False. Deal with it because the authoritative end-of-phase sign; don&#8217;t assume current_step == num_steps.<\/p>\n<p>Cancel latency is bounded however not zero. The builder finishes the present step earlier than checking the return worth. Lengthy tactic-search steps can push this into the seconds-to-tens-of-seconds vary.<\/p>\n<p>Thread security is required. The identical monitor occasion known as from a number of builder threads; uninstrumented dict or unordered_map entry from _render() will finally crash or tear.<\/p>\n<h2 id=\"get_started\" class=\"wp-block-heading\">Get began<\/h2>\n<p class=\"wp-block-paragraph\">The quickest technique to run this finish to finish is:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\ngit clone &#8211;depth 1 https:\/\/github.com\/NVIDIA\/TensorRT.git<br \/>\ncd TensorRT\/samples\/python\/simple_progress_monitor<br \/>\npython3 simple_progress_monitor.py\n<\/div>\n<p class=\"wp-block-paragraph\">This begins a dwell, animated construct of a ResNet-50 engine. Substitute simple_progress_monitor.py\u2018s monitor class with the model above or connect a cancel handler across the present class. C++ equal is on the market in samples\/sampleProgressMonitor\/.<\/p>\n<p class=\"wp-block-paragraph\">For bigger methods, the fitting subsequent step is changing the terminal renderer with the transport the appliance already makes use of comparable to Language Server Protocol notifications, server-sent occasions, or structured tool-call chunks. IProgressMonitor turns into the purpose the place TensorRT construct progress is translated into the appliance\u2019s progress mannequin.<\/p>\n<h3 id=\"learn_more\" class=\"wp-block-heading\">Study extra<\/h3>\n<p class=\"wp-block-paragraph\">Confer with the next assets for extra info:<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>A TensorRT engine construct can take seconds to many minutes. Giant strongly typed fashions, deep tactic search, and a chilly timing cache on a brand-new GPU SKU can depart builders, finish customers, or AI brokers looking at a frozen terminal with no concept whether or not to attend, retry, or kill the method. Most NVIDIA [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":2734,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[307,3211,1621,209,81,3210,219,2423],"class_list":["post-2732","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-builds","tag-cancelable","tag-engine","tag-longrunning","tag-nvidia","tag-observable","tag-python","tag-tensorrt"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++ - Future News 24<\/title>\n<meta name=\"description\" content=\"A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand&#x2d;new GPU SKU can&#8230;\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++ - Future News 24\" \/>\n<meta property=\"og:description\" content=\"A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand&#x2d;new GPU SKU can&#8230;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-22T16:35:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-23T09:59:16+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++\",\"datePublished\":\"2026-07-22T16:35:00+00:00\",\"dateModified\":\"2026-07-23T09:59:16+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/\"},\"wordCount\":2157,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp\",\"keywords\":[\"Builds\",\"Cancelable\",\"Engine\",\"LongRunning\",\"NVIDIA\",\"Observable\",\"Python\",\"TensorRT\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/\",\"name\":\"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++ - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp\",\"datePublished\":\"2026-07-22T16:35:00+00:00\",\"dateModified\":\"2026-07-23T09:59:16+00:00\",\"description\":\"A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand&#x2d;new GPU SKU can&#8230;\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/22\\\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++ - Future News 24","description":"A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand&#x2d;new GPU SKU can&#8230;","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/","og_locale":"en_US","og_type":"article","og_title":"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++ - Future News 24","og_description":"A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand&#x2d;new GPU SKU can&#8230;","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/","og_site_name":"Future News 24","article_published_time":"2026-07-22T16:35:00+00:00","article_modified_time":"2026-07-23T09:59:16+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++","datePublished":"2026-07-22T16:35:00+00:00","dateModified":"2026-07-23T09:59:16+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/"},"wordCount":2157,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp","keywords":["Builds","Cancelable","Engine","LongRunning","NVIDIA","Observable","Python","TensorRT"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/","name":"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++ - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp","datePublished":"2026-07-22T16:35:00+00:00","dateModified":"2026-07-23T09:59:16+00:00","description":"A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand&#x2d;new GPU SKU can&#8230;","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/NVIDIA-NCCL-Inspector-Real-Time-Performance-Monitoring.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/22\/make-long-running-nvidia-tensorrt-engine-builds-observable-and-cancelable-in-python-or-c\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Make Lengthy-Operating NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2732","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=2732"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2732\/revisions"}],"predecessor-version":[{"id":2733,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2732\/revisions\/2733"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/2734"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=2732"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=2732"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=2732"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}