{"id":4406,"date":"2026-08-28T17:06:00","date_gmt":"2026-08-28T17:06:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/"},"modified":"2026-08-29T15:59:11","modified_gmt":"2026-08-29T15:59:11","slug":"deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/","title":{"rendered":"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">Open AI fashions are evolving quicker than ever, however bringing them into native functions can nonetheless require model-specific conversion, preprocessing, post-processing, and runtime code.<\/p>\n<p class=\"wp-block-paragraph\">NVIDIA TensorRT Mannequin Join open assortment of reference implementations helps to handle this problem. TensorRT Mannequin Join exhibits you the best way to run supported fashions with NVIDIA TensorRT in native C++ functions. You need to use, examine, modify, and prolong the implementations. Mannequin Join is designed to help the open mannequin ecosystem wherever TensorRT runs.<\/p>\n<p class=\"wp-block-paragraph\">This publish explains what NVIDIA TensorRT Mannequin Join is and the best way to deploy a mannequin from Hugging Face mannequin ID to native C++ inference in two instructions. It additionally covers the 2 API ranges TensorRT Mannequin Join gives, the best way to combine customized GPU kernels, and the way the challenge is constructed to maintain tempo with the open mannequin ecosystem.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a9301acbfd59&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a9301acbfd59\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1999\" height=\"554\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck.webp\" alt=\"Diagram showing the out-of-framework deployment pipeline from PyTorch Model (research checkpoint) to ONNX or TorchScript (exchange format) to TensorRT Engine (optimized inference) to C++ Production (native runtime). Dashed arrows indicate fragile steps that block or delay deployment, with failure modes labeled at each transition: export failures and unsupported operations, operator gaps and accuracy re-validation, and custom C++ plugins and compounding maintenance.&#10;\" class=\"wp-image-121959\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-179x50.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-300x83.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-768x213.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-625x173.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-1536x426.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-645x179.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-500x139.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-160x44.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-362x100.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-397x110.png 397w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-1024x284.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-960x266.png 960w\" sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"554\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck.webp\" alt=\"Diagram showing the out-of-framework deployment pipeline from PyTorch Model (research checkpoint) to ONNX or TorchScript (exchange format) to TensorRT Engine (optimized inference) to C++ Production (native runtime). Dashed arrows indicate fragile steps that block or delay deployment, with failure modes labeled at each transition: export failures and unsupported operations, operator gaps and accuracy re-validation, and custom C++ plugins and compounding maintenance.&#10;\" class=\"lazyload wp-image-121959\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck.webp 1999w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-179x50.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-300x83.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-768x213.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-625x173.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-1536x426.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-645x179.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-500x139.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-160x44.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-362x100.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-397x110.png 397w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-1024x284.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/out-of-framework-model-deployment-bottleneck-960x266.png 960w\" data-sizes=\"(max-width: 1999px) 100vw, 1999px\"\/><figcaption class=\"wp-element-caption\">Determine 1. Out-of-framework deployment bottleneck<\/figcaption><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a9301acc0b1e&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a9301acc0b1e\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1318\" height=\"1088\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack.webp\" alt=\"Diagram showing the TensorRT Model Connect stack. Model weights from a Hugging Face or local checkpoint feed into TensorRT Model Connect, which contains model implementations across 80+ model families including Nemotron Speech and Qwen 3 VL, continuously extended by an Agentic Model Implementation Workflow. The model implementation and weights combine into a User Application layer, which is built on three stacked components: Model Connect Task APIs supporting Text, Vision, and Audio; TensorRT; and hardware targets including X86, ARM, DRIVE AGX, and Jetson AGX.&#10;\" class=\"wp-image-121963\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack.webp 1318w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-139x115.png 139w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-300x248.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-768x634.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-625x516.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-645x532.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-363x300.png 363w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-109x90.png 109w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-362x299.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-133x110.png 133w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-1024x845.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-654x540.png 654w\" sizes=\"(max-width: 1318px) 100vw, 1318px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1318\" height=\"1088\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack.webp\" alt=\"Diagram showing the TensorRT Model Connect stack. Model weights from a Hugging Face or local checkpoint feed into TensorRT Model Connect, which contains model implementations across 80+ model families including Nemotron Speech and Qwen 3 VL, continuously extended by an Agentic Model Implementation Workflow. The model implementation and weights combine into a User Application layer, which is built on three stacked components: Model Connect Task APIs supporting Text, Vision, and Audio; TensorRT; and hardware targets including X86, ARM, DRIVE AGX, and Jetson AGX.&#10;\" class=\"lazyload wp-image-121963\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack.webp 1318w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-139x115.png 139w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-300x248.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-768x634.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-625x516.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-645x532.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-363x300.png 363w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-109x90.png 109w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-362x299.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-133x110.png 133w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-1024x845.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/nvidia-tensorrt-model-connect-stack-654x540.png 654w\" data-sizes=\"(max-width: 1318px) 100vw, 1318px\"\/><figcaption class=\"wp-element-caption\">Determine 2. NVIDIA TensorRT Mannequin Join stack<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"how_to_deploy_a_model_from_model_id_to_native_c++_inference_in_two_commands\" class=\"wp-block-heading\">How one can deploy a mannequin from mannequin ID to native C++ inference in two instructions<\/h2>\n<p class=\"wp-block-paragraph\">Getting a mannequin into manufacturing mustn&#8217;t require deep compiler experience. Mannequin Join splits deployment into two phases with a single artifact between them.<\/p>\n<h3 id=\"1_build_the_bundle_python_cli\u00a0\" class=\"wp-block-heading\">1. Construct the bundle (Python CLI)\u00a0<\/h3>\n<p class=\"wp-block-paragraph\">For a supported mannequin, the primary section is constructing a deployment bundle from a Hugging Face mannequin ID or native checkpoint:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\ntrtmc construct Qwen\/Qwen3-0.6B -o qwen3-0.6B.bundle\n<\/div>\n<p class=\"wp-block-paragraph\">The bundle incorporates the TensorRT engines and the model-specific belongings wanted at runtime.\u00a0<\/p>\n<h3 id=\"2_load_and_run_c++\" class=\"wp-block-heading\">2. Load and run (C++)<\/h3>\n<p class=\"wp-block-paragraph\">Within the second section, a local C++ software then masses the bundle and works with task-level inputs and outputs:<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\n#embody<br \/>\nauto pipeline = trtmc::load(&#8220;qwen3-0.6b.bundle&#8221;);<br \/>\nauto consequence   = pipeline-&gt;generate(&#8220;Clarify why native inference issues.&#8221;, {.max_new_tokens = 20});<br \/>\nstd::cout &lt;&lt; consequence.textual content &lt;&lt; std::endl;\n<\/div>\n<p class=\"wp-block-paragraph\">Mannequin Join handles checkpoint mapping, TensorRT engine building, preprocessing, runtime orchestration, and post-processing. You begin with an entire working implementation as a substitute of rebuilding this integration for each mannequin household.<\/p>\n<p class=\"wp-block-paragraph\">You need to use Python to arrange the mannequin, however the deployed software runs natively with out requiring PyTorch or a Python interpreter in its manufacturing runtime.<\/p>\n<h2 id=\"two_api_levels_one_starting_point\" class=\"wp-block-heading\">Two API ranges, one place to begin<\/h2>\n<p class=\"wp-block-paragraph\">Mannequin Join gives two ranges of C++ APIs. With the semantic API, you possibly can work with acquainted inputs and outputs, resembling prompts, photos, and audio, whereas Mannequin Join handles model-specific preprocessing, execution, and post-processing.<\/p>\n<p class=\"wp-block-paragraph\">If you happen to want extra management, the module-level API permits you to work straight with named tensors and particular person TensorRT parts to customise the inference pipeline. Each APIs use the identical Mannequin Join implementations, so you can begin with a easy task-level interface and customise the pipeline solely when wanted.<\/p>\n<h2 id=\"extend_tensorrt_model_connect_with_custom_kernels\" class=\"wp-block-heading\">Lengthen TensorRT Mannequin Join with customized kernels<\/h2>\n<p class=\"wp-block-paragraph\">TVM FFI gives a language-agnostic interface for invoking GPU kernels with out tightly coupling the calling system to the kernel\u2019s implementation framework or runtime. Utilizing TVM FFI by means of TensorRT Mannequin Join, you possibly can substitute a focused portion of a mannequin with a customized GPU kernel whereas TensorRT continues to execute the remainder of the inference pipeline. This makes it simpler to combine specialised or newly developed kernels with out rebuilding the appliance round a separate runtime. See the Deliver Your Personal Kernel tutorial for a labored instance.<\/p>\n<h2 id=\"reference_implementations_for_the_open_model_ecosystem\" class=\"wp-block-heading\">Reference implementations for the open mannequin ecosystem<\/h2>\n<p class=\"wp-block-paragraph\">Mannequin Join will not be a brand new inference framework or a substitute for TensorRT. It&#8217;s a bridge between the end-to-end inference expertise for open fashions and the power of TensorRT to translate a computation graph into an accelerated engine on GPU.<\/p>\n<p class=\"wp-block-paragraph\">Every mannequin\u2019s implementation serves three functions:<\/p>\n<p>Operating a supported open mannequin in a local TensorRT-enabled software<\/p>\n<p>Studying from an entire, inspectable implementation of the mannequin and its inference pipeline<\/p>\n<p>Extending the implementation for a associated structure, customized checkpoint, or software requirement<\/p>\n<p class=\"wp-block-paragraph\">This gives the broader ecosystem with a clearer path to TensorRT deployment from a mannequin ID. Utility builders can start with working code. Group contributors can reuse present patterns so as to add help for brand spanking new fashions as a substitute of ranging from zero.<\/p>\n<p class=\"wp-block-paragraph\">The aim is easy: wherever TensorRT is out there, you need to have a constant Mannequin Join path for supported open fashions.<\/p>\n<h2 id=\"built_ai-natively_to_keep_pace_with_open_models\" class=\"wp-block-heading\">Constructed AI-natively to maintain tempo with open fashions<\/h2>\n<p class=\"wp-block-paragraph\">The open mannequin ecosystem adjustments shortly. New architectures and checkpoints seem repeatedly, so a reference library should evolve simply as shortly.<\/p>\n<p class=\"wp-block-paragraph\">Mannequin Join is constructed as an AI-native software program challenge. Coding brokers generate implementation code, checks, integrations, and documentation underneath human course and assessment. This allows the challenge to develop and validate a number of mannequin implementations in parallel whereas sustaining a constant structure and consumer expertise.<\/p>\n<p class=\"wp-block-paragraph\">Mannequin Join makes use of nightly releases to shorten the trail from a brand new mannequin, consumer report, or contribution to an accessible implementation. Automated validation stays the discharge gate.\u00a0 The quicker cadence helps new mannequin help, fixes, and UX enhancements attain you sooner.<\/p>\n<h2 id=\"delivering_the_complete_tensorrt_workflow\" class=\"wp-block-heading\">Delivering the whole TensorRT workflow<\/h2>\n<p class=\"wp-block-paragraph\">Mannequin Join is constructed on TensorRT, so efficiency stays central. For supported and validated workloads, Mannequin Join can ship quicker inference than torch.compile, and every implementation is repeatedly examined and optimized because the challenge evolves.<\/p>\n<p class=\"wp-block-paragraph\">Efficiency mustn&#8217;t come on the expense of usability. Mannequin Join brings the whole workflow collectively: discover the mannequin ID, construct the mannequin, load it from C++, and adapt it when wanted. You get an accessible path to high-performance TensorRT inference whereas retaining the power to examine, customise, and optimize the underlying inference pipeline.<\/p>\n<h2 id=\"get_started_with_nvidia_tensorrt_model_connect\" class=\"wp-block-heading\">Get began with NVIDIA TensorRT Mannequin Join<\/h2>\n<p class=\"wp-block-paragraph\">Go to the NVIDIA\/TensorRT-Mannequin-Join GitHub repo to search out supported implementations and construct a mannequin bundle. Use an implementation as-is, adapt it in your software, or contribute help that helps the following developer carry one other open mannequin to TensorRT.<\/p>\n<p class=\"wp-block-paragraph\">Need to use an AI-native fast begin that doesn\u2019t require an advanced setup? Open a terminal in any folder you possibly can entry, then paste the next immediate right into a coding agent. You need to have an entire deployment in minutes.\u00a0<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\n\/aim Clone https:\/\/github.com\/NVIDIA\/TensorRT-Mannequin-Join.git into<br \/>\na brand new TensorRT-Mannequin-Join listing within the present workspace. Detect<br \/>\nthe present GPU compute functionality, modify the repository improvement Docker<br \/>\npicture, construct and begin the container, set up TensorRT-Mannequin-Join, compile<br \/>\nthe CLI, TensorRT backend, and all native mannequin DSOs just for that SM, then<br \/>\nconstruct and run an end-to-end Qwen\/Qwen3-0.6B smoke check. Don&#8217;t commit or push<br \/>\nadjustments. Report the results of the check, present precise command, enter and output of<br \/>\nthe inference run.\n<\/div>\n<p class=\"wp-block-paragraph\">For extra about mannequin protection, structure particulars, and a full developer information, see the TensorRT Mannequin Join documentation.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Open AI fashions are evolving quicker than ever, however bringing them into native functions can nonetheless require model-specific conversion, preprocessing, post-processing, and runtime code. NVIDIA TensorRT Mannequin Join open assortment of reference implementations helps to handle this problem. TensorRT Mannequin Join exhibits you the best way to run supported fashions with NVIDIA TensorRT in native [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":4408,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/ai-use-cases-1.webp","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[2077,1469,2959,489,1068,105,81,35,2423],"class_list":["post-4406","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-checkpoint","tag-commands","tag-connect","tag-deploy","tag-inference","tag-model","tag-nvidia","tag-open","tag-tensorrt"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join - Future News 24<\/title>\n<meta name=\"description\" content=\"Open AI models are evolving faster than ever, but bringing them into native applications can still require model&#x2d;specific conversion, preprocessing&#8230;\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Open AI models are evolving faster than ever, but bringing them into native applications can still require model&#x2d;specific conversion, preprocessing&#8230;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-28T17:06:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-29T15:59:11+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/ai-use-cases-1.webp\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/ai-use-cases-1.webp\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"6 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join\",\"datePublished\":\"2026-08-28T17:06:00+00:00\",\"dateModified\":\"2026-08-29T15:59:11+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/\"},\"wordCount\":1138,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/ai-use-cases-1.webp\",\"keywords\":[\"Checkpoint\",\"commands\",\"Connect\",\"Deploy\",\"inference\",\"Model\",\"NVIDIA\",\"Open\",\"TensorRT\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/\",\"name\":\"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/ai-use-cases-1.webp\",\"datePublished\":\"2026-08-28T17:06:00+00:00\",\"dateModified\":\"2026-08-29T15:59:11+00:00\",\"description\":\"Open AI models are evolving faster than ever, but bringing them into native applications can still require model&#x2d;specific conversion, preprocessing&#8230;\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/ai-use-cases-1.webp\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/ai-use-cases-1.webp\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/28\\\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join - Future News 24","description":"Open AI models are evolving faster than ever, but bringing them into native applications can still require model&#x2d;specific conversion, preprocessing&#8230;","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/","og_locale":"en_US","og_type":"article","og_title":"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join - Future News 24","og_description":"Open AI models are evolving faster than ever, but bringing them into native applications can still require model&#x2d;specific conversion, preprocessing&#8230;","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/","og_site_name":"Future News 24","article_published_time":"2026-08-28T17:06:00+00:00","article_modified_time":"2026-08-29T15:59:11+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/ai-use-cases-1.webp","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/ai-use-cases-1.webp","twitter_misc":{"Written by":"Future News 24","Est. reading time":"6 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join","datePublished":"2026-08-28T17:06:00+00:00","dateModified":"2026-08-29T15:59:11+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/"},"wordCount":1138,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/ai-use-cases-1.webp","keywords":["Checkpoint","commands","Connect","Deploy","inference","Model","NVIDIA","Open","TensorRT"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/","name":"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/ai-use-cases-1.webp","datePublished":"2026-08-28T17:06:00+00:00","dateModified":"2026-08-29T15:59:11+00:00","description":"Open AI models are evolving faster than ever, but bringing them into native applications can still require model&#x2d;specific conversion, preprocessing&#8230;","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/ai-use-cases-1.webp","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/08\/ai-use-cases-1.webp"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/28\/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Deploy an Open Mannequin from Checkpoint to Inference in Two Instructions with NVIDIA TensorRT Mannequin Join"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4406","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=4406"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4406\/revisions"}],"predecessor-version":[{"id":4407,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/4406\/revisions\/4407"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/4408"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=4406"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=4406"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=4406"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}