{"id":656,"date":"2026-06-01T04:43:00","date_gmt":"2026-06-01T04:43:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/"},"modified":"2026-06-07T10:59:32","modified_gmt":"2026-06-07T10:59:32","slug":"develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/","title":{"rendered":"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p>Bodily AI programs should perceive the actual world earlier than they will act inside it. Robots, autonomous autos, and sensible areas want to grasp what\u2019s taking place of their world, predict what\u2019s prone to occur subsequent, and generate actions for particular environments, embodiments, and duties.<\/p>\n<p>NVIDIA Cosmos 3 is a frontier basis mannequin for bodily AI that mixes bodily reasoning, world era, and motion era inside a single open mannequin.\u00a0<\/p>\n<p>NVIDIA is open sourcing Cosmos 3 fashions, coaching scripts, deployment instruments, and datasets to make bodily AI improvement extra open and reproducible. This weblog submit covers the basics of Cosmos 3, highlights key ideas from the technical report, guides by means of technical workflows, and exhibits how groups constructing robotic manipulation programs, autonomous autos, and warehouse monitoring options can get began.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a254eff7dd1b&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a254eff7dd1b\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"512\" height=\"288\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light.gif\" alt=\"A video clip generated by Cosmos 3 for the autonomous driving domain. The video is from a vehicle\u2019s point-of-view at an intersection. Another car crosses the intersection in front of this vehicle, and then the vehicle takes a left turn. The video looks realistic and shows houses, trees, and cars in the surroundings.\" class=\"wp-image-117482\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light.gif 512w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-179x101.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-300x169.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-500x281.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-160x90.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-362x204.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-196x110.gif 196w\" sizes=\"(max-width: 512px) 100vw, 512px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"512\" height=\"288\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light.gif\" alt=\"A video clip generated by Cosmos 3 for the autonomous driving domain. The video is from a vehicle\u2019s point-of-view at an intersection. Another car crosses the intersection in front of this vehicle, and then the vehicle takes a left turn. The video looks realistic and shows houses, trees, and cars in the surroundings.\" class=\"lazyload wp-image-117482\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light.gif 512w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-179x101.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-300x169.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-500x281.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-160x90.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-362x204.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Traffic-Light-196x110.gif 196w\" data-sizes=\"(max-width: 512px) 100vw, 512px\"\/><figcaption class=\"wp-element-caption\">Determine 1. A clip of a video generated by Cosmos 3 for the autonomous driving area<\/figcaption><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a254eff7e615&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a254eff7e615\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"832\" height=\"480\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire.gif\" alt=\"A video shows a corridor with shelves of boxes on either side and a pile of boxes on the ground. Three people are standing next to the pile of boxes. There\u2019s a small explosion from one of the boxes on the floor, and it starts smoking.\u00a0\" class=\"wp-image-117618\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire.gif 832w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-179x103.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-300x173.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-768x443.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-625x361.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-645x372.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-500x288.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-156x90.gif 156w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-362x209.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-191x110.gif 191w\" sizes=\"(max-width: 832px) 100vw, 832px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"832\" height=\"480\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire.gif\" alt=\"A video shows a corridor with shelves of boxes on either side and a pile of boxes on the ground. Three people are standing next to the pile of boxes. There\u2019s a small explosion from one of the boxes on the floor, and it starts smoking.\u00a0\" class=\"lazyload wp-image-117618\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire.gif 832w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-179x103.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-300x173.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-768x443.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-625x361.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-645x372.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-500x288.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-156x90.gif 156w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-362x209.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Fire-191x110.gif 191w\" data-sizes=\"(max-width: 832px) 100vw, 832px\"\/><figcaption class=\"wp-element-caption\">Determine 2. A video generated utilizing Cosmos 3 for warehouse security information.<\/figcaption><\/figure>\n<\/div>\n<p>Key highlights of this launch embody:<\/p>\n<p>NVIDIA Cosmos 3 Nano and NVIDIA Cosmos 3 Tremendous mannequin checkpoints on Hugging Face with code on GitHub.<\/p>\n<p>Open datasets for bodily AI purposes like robotics and autonomous driving.<\/p>\n<p>Open post-training scripts for adapting Cosmos 3 to your area.<\/p>\n<p>Cosmos NIM microservices for simple, optimized deployment on NVIDIA GPUs.<\/p>\n<h2 id=\"what\u2019s_new_in_cosmos_3\" class=\"wp-block-heading\">What\u2019s new in Cosmos 3<\/h2>\n<p>Earlier Cosmos releases separated world era, bodily understanding, and managed scene era into completely different fashions and workflows. This launch unifies these capabilities with a Combination-of-Transformers (MoT) structure constructed round two towers.\u00a0<\/p>\n<p>Reasoner tower: A vision-language mannequin (VLM) that interprets multimodal observations like photographs, movies, and textual content. This tower makes use of an autoregressive structure to interpret the enter and perceive movement, object interactions, and different bodily context. This serves because the \u2018mind\u2019 that causes concerning the world earlier than any era occurs.<\/p>\n<p>Generator tower: Generates future observations and motion sequences. This tower makes use of a diffusion-based course of to generate physics-aware video and motion outputs which might be conditioned on the reasoner tower\u2019s understanding. The reasoner might be known as independently, however the generator at all times prompts each towers for guided era.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a254eff7f4fa&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a254eff7f4fa\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"1680\" height=\"694\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151.webp\" alt=\"Cosmos 3 architecture diagram: an autoregressive reasoner tower that takes in text, image, video, audio, and action inputs is connected to a diffusion-based generator tower that outputs text, image, video, audio, and action. Information from the reasoner tower feeds unidirectionally into the generator tower, which enables coherent generation.\" class=\"wp-image-117483\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151.webp 1680w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-179x74.webp 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-300x124.webp 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-768x317.webp 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-625x258.webp 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-1536x635.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-645x266.webp 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-500x207.webp 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-160x66.webp 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-362x150.webp 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-266x110.webp 266w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-1024x423.webp 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-960x397.webp 960w\" sizes=\"(max-width: 1680px) 100vw, 1680px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1680\" height=\"694\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151.webp\" alt=\"Cosmos 3 architecture diagram: an autoregressive reasoner tower that takes in text, image, video, audio, and action inputs is connected to a diffusion-based generator tower that outputs text, image, video, audio, and action. Information from the reasoner tower feeds unidirectionally into the generator tower, which enables coherent generation.\" class=\"lazyload wp-image-117483\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151.webp 1680w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-179x74.webp 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-300x124.webp 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-768x317.webp 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-625x258.webp 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-1536x635.webp 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-645x266.webp 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-500x207.webp 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-160x66.webp 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-362x150.webp 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-266x110.webp 266w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-1024x423.webp 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/image-11-e1780000686151-960x397.webp 960w\" data-sizes=\"(max-width: 1680px) 100vw, 1680px\"\/><figcaption class=\"wp-element-caption\">Determine 3. Cosmos 3 structure\u00a0<\/figcaption><\/figure>\n<\/div>\n<p>This structure allows a single mannequin to do reasoning and era duties, simplifying improvement by eliminating orchestration between a number of fashions and inference pipelines.\u00a0<\/p>\n<h3 id=\"choose_the_right_model_size\" class=\"wp-block-heading\">Select the proper mannequin dimension<\/h3>\n<p>Two Cosmos 3 fashions are at present out there:<\/p>\n<p>Cosmos 3 Nano is the compact model with 16B parameters and optimized for environment friendly inference. It\u2019s designed to run on workstation-grade compute, just like the NVIDIA RTX PRO 6000 GPU for real-time robotics inference and bodily AI purposes.<\/p>\n<p>Cosmos 3 Tremendous is a 64B parameter mannequin designed for max high quality and functionality. It delivers the best benchmark scores and targets datacenter deployment on NVIDIA Hopper and NVIDIA Blackwell GPUs, making it appropriate for large-scale artificial information era and superior bodily reasoning workloads.\u00a0<\/p>\n<h3 id=\"supported_modalities\" class=\"wp-block-heading\">Supported modalities<\/h3>\n<p>Cosmos 3 helps the next enter and output modalities by means of its unified structure:<\/p>\n<figure class=\"wp-block-table aligncenter\">InputOutputApplicationTextImagePhysically-plausible Picture generationText | VideoVideoWorld mannequin for uncommon edge case video information generationText | ImageVideoWorld mannequin for predictionText | Picture | VideoTextVLM for reasoningAction | Video | TextVideoAction-conditioned world modelVideo | TextVideo | ActionWorld motion mannequin, video motion mannequin, imaginative and prescient language motion mannequin,\u00a0coverage mannequin for robotic studying\u00a0<figcaption class=\"wp-element-caption\">Desk 1. Enter and output modalities supported by Cosmos 3 for various purposes<\/figcaption><\/figure>\n<h3 id=\"open_datasets_for_physical_ai\" class=\"wp-block-heading\">Open datasets for bodily AI<\/h3>\n<p>With the Cosmos 3 launch, NVIDIA is open-sourcing six artificial information era (SDG) datasets on Hugging Face. These cowl robotics, physics simulation, spatial reasoning, human movement, driving, and warehouse environments, and can be utilized for post-training Cosmos 3 and different fashions:<\/p>\n<p>Bodily AI World Mannequin Artificial Datasets embody:<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a254eff8043e&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a254eff8043e\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"720\" height=\"405\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot.gif\" alt=\"A collection of videos in the Embodied Robot Scenes dataset. The videos show different humanoid robots doing manipulation tasks in different environments.\" class=\"wp-image-117621\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot.gif 720w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-179x101.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-300x169.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-625x352.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-645x363.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-660x370.gif 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-500x281.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-160x90.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-362x204.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-196x110.gif 196w\" sizes=\"(max-width: 720px) 100vw, 720px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"720\" height=\"405\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot.gif\" alt=\"A collection of videos in the Embodied Robot Scenes dataset. The videos show different humanoid robots doing manipulation tasks in different environments.\" class=\"lazyload wp-image-117621\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot.gif 720w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-179x101.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-300x169.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-625x352.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-645x363.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-660x370.gif 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-500x281.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-160x90.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-362x204.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Embodied-Robot-196x110.gif 196w\" data-sizes=\"(max-width: 720px) 100vw, 720px\"\/><figcaption class=\"wp-element-caption\">Determine 4. Manipulation examples from the Embodied Robotic Scenes dataset\u00a0<\/figcaption><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a254eff80e48&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a254eff80e48\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"1065\" height=\"624\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1.gif\" alt=\"A collection of videos in the Physical Interaction Scenes dataset. The videos show simulated scenes like a wrecking ball hitting objects, a toy tower collapsing, and dominoes falling. For each scene, the dataset has corresponding ground-truth physics annotations like per-object velocity, center-of-mass displacement, and per-frame semantic segmentation.\" class=\"wp-image-117625\" style=\"aspect-ratio:1.7067750333354499;width:720px;height:auto\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1.gif 1065w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-179x105.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-300x176.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-768x450.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-625x366.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-645x378.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-500x293.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-154x90.gif 154w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-362x212.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-188x110.gif 188w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-1024x600.gif 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-922x540.gif 922w\" sizes=\"(max-width: 1065px) 100vw, 1065px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1065\" height=\"624\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1.gif\" alt=\"A collection of videos in the Physical Interaction Scenes dataset. The videos show simulated scenes like a wrecking ball hitting objects, a toy tower collapsing, and dominoes falling. For each scene, the dataset has corresponding ground-truth physics annotations like per-object velocity, center-of-mass displacement, and per-frame semantic segmentation.\" class=\"lazyload wp-image-117625\" style=\"aspect-ratio:1.7067750333354499;width:720px;height:auto\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1.gif 1065w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-179x105.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-300x176.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-768x450.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-625x366.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-645x378.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-500x293.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-154x90.gif 154w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-362x212.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-188x110.gif 188w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-1024x600.gif 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Physical-Scene-1-922x540.gif 922w\" data-sizes=\"(max-width: 1065px) 100vw, 1065px\"\/><figcaption class=\"wp-element-caption\">Determine 5. Examples from the Bodily Interplay Scenes dataset<\/figcaption><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a254eff81997&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a254eff81997\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"1760\" height=\"894\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning.webp\" alt=\"A collection of images showing the Spatial Reasoning dataset, including scenes like kitchens, corridors, offices, and utility rooms. It also includes question-answer pairs like, \u201cHow far is the coffee table from the sofa?\u201d and \u201cWhat is the best route for the robot to reach the study room?\u201d\u00a0\" class=\"wp-image-117630\" style=\"width:719px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning.webp 1760w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-179x91.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-300x152.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-768x390.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-625x317.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-1536x780.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-645x328.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-500x254.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-160x81.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-362x184.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-217x110.png 217w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-1024x520.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-960x488.png 960w\" sizes=\"(max-width: 1760px) 100vw, 1760px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1760\" height=\"894\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning.webp\" alt=\"A collection of images showing the Spatial Reasoning dataset, including scenes like kitchens, corridors, offices, and utility rooms. It also includes question-answer pairs like, \u201cHow far is the coffee table from the sofa?\u201d and \u201cWhat is the best route for the robot to reach the study room?\u201d\u00a0\" class=\"lazyload wp-image-117630\" style=\"width:719px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning.webp 1760w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-179x91.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-300x152.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-768x390.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-625x317.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-1536x780.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-645x328.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-500x254.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-160x81.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-362x184.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-217x110.png 217w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-1024x520.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Spatial-Reasoning-960x488.png 960w\" data-sizes=\"(max-width: 1760px) 100vw, 1760px\"\/><figcaption class=\"wp-element-caption\">Determine 6. Examples from the Spatial Reasoning dataset<\/figcaption><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a254eff82335&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a254eff82335\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"1280\" height=\"720\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset.gif\" alt=\"A collection of videos in the Digital Human Scenes dataset. The videos show some simulated indoor and outdoor environments with digital people standing and moving. These videos provide diverse human appearance, motion, scene context, lighting, and camera motion.\" class=\"wp-image-117485\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset.gif 1280w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-179x101.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-300x169.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-768x432.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-625x352.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-645x363.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-660x370.gif 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-500x281.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-160x90.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-362x204.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-196x110.gif 196w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-1024x576.gif 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-960x540.gif 960w\" sizes=\"(max-width: 1280px) 100vw, 1280px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1280\" height=\"720\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset.gif\" alt=\"A collection of videos in the Digital Human Scenes dataset. The videos show some simulated indoor and outdoor environments with digital people standing and moving. These videos provide diverse human appearance, motion, scene context, lighting, and camera motion.\" class=\"lazyload wp-image-117485\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset.gif 1280w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-179x101.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-300x169.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-768x432.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-625x352.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-645x363.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-660x370.gif 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-500x281.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-160x90.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-362x204.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-196x110.gif 196w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-1024x576.gif 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Human-Dataset-960x540.gif 960w\" data-sizes=\"(max-width: 1280px) 100vw, 1280px\"\/><figcaption class=\"wp-element-caption\">Determine 7. Examples from the Digital Human Scenes dataset<\/figcaption><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a254eff82c31&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a254eff82c31\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"960\" height=\"360\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving.gif\" alt=\"A collection of videos from the Autonomous Driving Scenarios dataset. The videos are from the ego point of view of an autonomous vehicle and show the vehicle driving on roads in different scenarios. The videos show diverse weather and lighting conditions and driving behaviors like lane changing and pedestrian interactions.\" class=\"wp-image-117487\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving.gif 960w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-179x67.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-300x113.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-768x288.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-625x234.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-645x242.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-500x188.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-160x60.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-362x136.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-293x110.gif 293w\" sizes=\"(max-width: 960px) 100vw, 960px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"960\" height=\"360\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving.gif\" alt=\"A collection of videos from the Autonomous Driving Scenarios dataset. The videos are from the ego point of view of an autonomous vehicle and show the vehicle driving on roads in different scenarios. The videos show diverse weather and lighting conditions and driving behaviors like lane changing and pedestrian interactions.\" class=\"lazyload wp-image-117487\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving.gif 960w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-179x67.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-300x113.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-768x288.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-625x234.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-645x242.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-500x188.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-160x60.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-362x136.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Self-Driving-293x110.gif 293w\" data-sizes=\"(max-width: 960px) 100vw, 960px\"\/><figcaption class=\"wp-element-caption\">Determine 8. Examples from the Autonomous Driving Eventualities dataset<\/figcaption><\/figure>\n<\/div>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a254eff83504&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a254eff83504\" class=\"aligncenter size-full is-resized wp-lightbox-container\"><img decoding=\"async\" width=\"768\" height=\"432\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations.gif\" alt=\"A collection of videos from the Warehouse Operations Scenes dataset. The videos show simulated warehouse scenes from different camera angles. Some videos show a forklift moving and colliding with people or objects. In another video, a person drops a cardboard box on the floor.\u00a0\u00a0\" class=\"wp-image-117488\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-179x101.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-300x169.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-625x352.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-645x363.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-660x370.gif 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-500x281.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-160x90.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-362x204.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-196x110.gif 196w\" sizes=\"(max-width: 768px) 100vw, 768px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"768\" height=\"432\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations.gif\" alt=\"A collection of videos from the Warehouse Operations Scenes dataset. The videos show simulated warehouse scenes from different camera angles. Some videos show a forklift moving and colliding with people or objects. In another video, a person drops a cardboard box on the floor.\u00a0\u00a0\" class=\"lazyload wp-image-117488\" style=\"width:720px\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations.gif 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-179x101.gif 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-300x169.gif 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-625x352.gif 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-645x363.gif 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-660x370.gif 660w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-500x281.gif 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-160x90.gif 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-362x204.gif 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Warehouse-Operations-196x110.gif 196w\" data-sizes=\"(max-width: 768px) 100vw, 768px\"\/><figcaption class=\"wp-element-caption\">Determine 9. Examples from the Warehouse Operations Scenes dataset<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"nvidia_cosmos_human_evaluation_benchmark\" class=\"wp-block-heading\">NVIDIA Cosmos Human Analysis benchmark<\/h2>\n<p>The NVIDIA Cosmos Human Analysis (HUE) framework assesses Cosmos 3 generator high quality throughout consultant area duties.<\/p>\n<p>As SOTA video era fashions saturate present automated leaderboards, rating variations between releases are sometimes too slender for significant comparability. HUE shifts analysis from subjective grading to goal truth verification, enabling fine-grained comparability between top-tier fashions. The result&#8217;s a extra dependable high quality sign for each fast iteration and rigorous launch choices backed by full human analysis.<\/p>\n<p>HUE evaluates video era high quality utilizing atomic binary verification. Every generated video is decomposed into single-fact sure\/no questions throughout 4 dimensions\u2014semantic alignment, bodily legal guidelines, geometric reasoning, and visible integrity\u2014spanning seven Bodily AI domains, together with robotics, autonomous autos, and physics. These questions are generated by a VLM pipeline, refined by human consultants, and launched as open supply on Hugging Face.<\/p>\n<h2 id=\"benchmark_results\" class=\"wp-block-heading\">Benchmark outcomes<\/h2>\n<p>Cosmos 3 has been evaluated throughout a number of benchmark suites masking bodily AI reasoning, era high quality, and domain-specific efficiency.<\/p>\n<p>Reasoning benchmarks<\/p>\n<p>Cosmos 3 Tremendous and Cosmos 3 Nano lead on VANTAGE-Bench on the 32B tier and the 8B tier, respectively:<\/p>\n<p>VANTAGE-Bench: \u00a0First public benchmark for evaluating vision-language fashions on real-world fixed-camera footage throughout warehouses, transportation, and sensible areas.\u00a0<\/p>\n<p>Site visitors Anomaly Reasoning (TAR): A brand new leaderboard for detecting and reasoning anomalous occasions in transportation footage and the official leaderboard for AI Metropolis Problem 2026 Observe 3.<\/p>\n<p>Generator benchmarks<\/p>\n<p>Cosmos 3 is the open-source SOTA and at present leads on PAI-Bench, R-Bench Physics-IQ, and RoboLab throughout public leaderboards:<\/p>\n<p>Synthetic Evaluation: A benchmarking platform that ranks AI fashions for textual content, picture, and video era. Cosmos 3 is the main open supply mannequin on the Textual content to Picture leaderboard and Picture to Video (no audio) leaderboard.<\/p>\n<p>R-Bench: A benchmark for evaluating video-based world fashions in robotic video era. It assesses job completion and visible high quality by means of sub-metrics like structural consistency, bodily plausibility, and execution completeness.<\/p>\n<p>PAI-Bench: A unified benchmark evaluating bodily AI throughout video understanding and video era, spanning domains like robotics, autonomous autos, and physics widespread sense.<\/p>\n<p>Physics-IQ: A benchmark of real-world movies that checks whether or not generative video fashions really perceive bodily rules, moderately than simply reaching visible realism.<\/p>\n<p>RoboLab: A simulation benchmark for evaluating task-generalist robotic insurance policies.<\/p>\n<h2 id=\"training_recipes\" class=\"wp-block-heading\">Coaching recipes<\/h2>\n<p>A central part of the Cosmos 3 launch is a completely open set of coaching recipes. Past mannequin checkpoints, this launch gives code, configs, and workflows for adapting Cosmos 3 to new domains, embodiments, and datasets.<\/p>\n<p>Supervised Tremendous-Tuning post-training<\/p>\n<p>Supervised Tremendous-Tuning (SFT) allows builders to adapt a Cosmos 3 mannequin to their very own information. The launched recipes embody imaginative and prescient era post-training for customized video datasets, in addition to action-oriented recipes for robotics and bodily AI workflows. Builders can customise Cosmos 3 for his or her goal domains throughout robotics, autonomous driving, and warehouse automation.<\/p>\n<p>The post-training code and configs can be found on GitHub.<\/p>\n<p>Motion post-training<\/p>\n<p>Motion post-training adapts Cosmos 3 for action-aware Bodily AI purposes, together with ahead dynamics, inverse dynamics, and coverage era. Builders can post-train Cosmos 3 on action-labeled information. For robotics purposes, this consists of a number of necessary workflows: producing future observations conditioned on robotic actions, inferring the actions behind noticed demonstrations, and predicting motion sequences from present observations and job prompts. This makes Cosmos 3 a powerful basis for world motion modeling and coverage studying.<\/p>\n<figure class=\"wp-block-embed aligncenter is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio\">\n<p>\n<span class=\"embed-youtube\" style=\"text-align:center; display: block;\"><\/span>\n<\/p><figcaption class=\"wp-element-caption\">Video 1. Tutorial video displaying  post-train Cosmos 3<\/figcaption><\/figure>\n<h2 id=\"deploy_with_nvidia_nim_microservices\" class=\"wp-block-heading\">Deploy with NVIDIA NIM Microservices<\/h2>\n<p>Cosmos 3 fashions are additionally out there as NVIDIA NIM microservices for optimized, production-ready deployment. NIM microservices bundle the mannequin with optimized inference runtimes, delivering excessive efficiency with out the necessity to manually tune serving infrastructure. NIM microservices are simpler to make use of for inference workflows in comparison with the Cosmos 3 repo on GitHub, which is most well-liked for post-training workflows.<\/p>\n<p>The Cosmos 3 Reasoner NIM is on the market at present, delivering the reasoning capabilities of the Cosmos 3 mannequin. Hold posted for the Cosmos 3 Generator NIM, which gives full era capabilities of the Cosmos 3 mannequin.<\/p>\n<p>Optimizations made to speed up inference<\/p>\n<p>Quantization: Cosmos 3 NIM helps choosing BF16, FP8, or NVFP4 quantized checkpoints. The NVFP4 quantization reduces the mannequin\u2019s numerical precision from BF16 to 4-bit floating level, reaching as much as 2x inference speedup.\u00a0<\/p>\n<p>vLLM: Is an open supply inference engine that makes use of strategies like steady batching, paged consideration, and tensor parallelism to serve LLMs effectively. The Cosmos 3 Reasoner NIM serving stack is constructed on vLLM for increased throughput in comparison with standard serving approaches. Cosmos 3 Nano is able to run with vLLM-omni and NVIDIA Dynamo for prime efficiency.<\/p>\n<p>Environment friendly Video Sampling (EVS): This system reduces the variety of video tokens fed into the VLM throughout inference, rushing up the Cosmos Purpose NIM. EVS works on the chunk stage, conserving essentially the most distinctive chunks of every body and pruning the remainder. Smaller GPUs have a tendency to learn extra from this system.<\/p>\n<p>Easy methods to run the NIM\u00a0<\/p>\n<p>An NVIDIA NGC API secret is required to drag the containers and obtain the Cosmos 3 fashions from NGC.<\/p>\n<p>To drag and run the Cosmos 3 Nano Reasoner NIM. For the Cosmos 3 Tremendous Reasoner NIM, specify NIM_MODEL_SIZE=tremendous.<\/p>\n<div class=\"wp-block-syntaxhighlighter-code \">\ndocker run &#8211;gpus=all<br \/>\n  -e NGC_API_KEY=$NGC_API_KEY<br \/>\n  -e NIM_MODEL_SIZE=nano<br \/>\n  -p 8000:8000<br \/>\n  nvcr.io\/nim\/nvidia\/cosmos3-reasoner:newest\n<\/div>\n<p>Discover particulars on API utilization and extra within the documentation.<\/p>\n<figure class=\"wp-block-embed aligncenter is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio\">\n<p>\n<span class=\"embed-youtube\" style=\"text-align:center; display: block;\"><\/span>\n<\/p><figcaption class=\"wp-element-caption\">Video 2. Tutorial video displaying  use the Cosmos Reasoner NIM<\/figcaption><\/figure>\n<h2 id=\"get_started\" class=\"wp-block-heading\">Get began<\/h2>\n<p>Acknowledgments<\/p>\n<p class=\"has-small-font-size\">Cosmos 3 is the results of wonderful collaboration between many groups and folks throughout NVIDIA, together with Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, Aarti Basant, Mukesh Beladiya, Mohammad Qazim Bhat, Zaid Pervaiz Bhat, Dan Blick, Vanni Brighella, Han Cai, Tiffany Cai, Eric Cameracci, Jiaxin Cao, Yulong Cao, Mark Carlson, Carlos Casanova, Ting-Yun Chang, Yan Chang, Yu-Wei Chao, Prithvijit Chattopadhyay, Roshan Chaudhari, Chieh-Yun Chen, Junyu Chen, Ke Chen, Qizhi Chen, Wenkai Chen, Xiaotong Chen, Yu Chen, An-Chieh Cheng, Click on Cheng, Xiu Chia, Jeana Choi, Chaeyeon Chung, Wenyan Cong, Yin Cui, Magdalena Dadela, Nalin Dadhich, Wenliang Dai, Joyjit Daw, Alperen Degirmenci, Rodrigo Vieira Del Monte, Robert Denomme, Sameer Dharur, Marco Di Lucca, Ke Ding, Wenhao Ding, Yifan Ding, Yuzhu Dong, Nicole Drumheller, Yilun Du, Aigul Dzhumamuratova, Aleksandr Efitorov, Hamid Eghbalzadeh, Naomi Eigbe, Imad El Hanafi, Hassan Eslami, Benedikt Falk, Jiaojiao Fan, Jim Fan, Amol Fasale, Sergiy Fefilatyev, Liang Feng, Francesco Ferroni, Sanja Fidler, Xiao Fu, Vikram Fugro, Prashant Gaikwad, TJ Galda, Katelyn Gao, Yihuai Gao, Wenhang Ge, Sreyan Ghosh, Arushi Goel, Vivek Goel, Akash Gokul, Rama Govindaraju, Jinwei Gu, Miguel Guerrero, Elfie Guo, Aryaman Gupta, Siddharth Gururani, Hugo Hadfield, Tune Han, Ankur Handa, Zekun Hao, Mohammad Harrim, Ali Hassani, Nathan Hayes-Roth, Yufan He, Chris Helvig, Cyrus Hogg, Madison Huang, Michael Huang, Sophia Huang, Yufan Huang, Jacob Huffman, DeLesley Hutchins, Suneel Indupuru, Boris Ivanovic, Arihant Jain, Joel Jang, Ryan Ji, Yanan Jian, Dongfu Jiang, Jingyi Jin, Atharva Joshi, Nikhilesh Joshi, Pranjali Joshi, Jaehun Jung, Weiwei Kang, Scott Kassekert, Jan Kautz, Ashna Khetan, Julia Kiczka, Slawek Kierat, Gwanghyun Kim, Kuno Kim, Sunny Kim, Kezhi Kong, Xin Kong, Zhifeng Kong, Tomasz Kornuta, Egor Krivov, Hui Kuang, Saurav Kumar, Chia-Wen Kuo, George Kurian, Wojciech Kutak, JF Lafleche, Himangshu Lahkar, Omar Laymoun, Jayjun Lee, Sanggil Lee, Gabriele Leone, Boyi Li, Freya Li, Jiajun Li, Jinfeng Li, Ling Li, Pengcheng Li, Shangru Li, Tingle Li, Xiaolong Li, Xuan Li, Zhaoshuo Li, Zhiqi Li, Hao Liang, Maosheng Liao, Chen-Hsuan Lin, Tsung-Yi Lin, Ming-Yu Liu, Sifei Liu, Zihan Liu, Hai Loc Lu, Xiangyu Lu, Alice Luo, Ruipu Luo, Wenjie Luo, Jiangran Lyu, Martin Ding Ma, Nic Ma, Qianli Ma, Dawid Majchrowski, Louis Marcoux, Miguel Martin, Qing Miao, Ashkan Mirzaei, Shreyas Misra, Kaichun Mo, Durra Mohsin, Hyejin Moon, Pawel Morkisz, Saeid Motiian, Kirill Motkov, Seungjun Nah, Yashraj Narang, Deepak Narayanan, Thabang Ngazimbi, Julian Ouyang, David Web page, Yatian Pang, Sehwi Park, Mahesh Patekar, Mostofa Patwary, Marco Pavone, Trung Pham, Wei Ping, Soha Pouya, Shrimai Prabhumoye, Varun Praveen, Delin Qu, Hesam Rabeti, Morteza Ramezanali, Marilyn Reeb, Xuanchi Ren, Kristen Rumley, Wojciech Rymer, Jun Saito, Yeongho Seol, John Shao, Piyush Shekdar, Tianwei Shen, Humphrey Shi, Min Shi, Stella Shi, Kevin Shih, Mohammad Shoeybi, Mateusz Sieniawski, Shuran Tune, Alexander Sotelo, Amir Sotoodeh, Sunil Srinivasa, Vignesh Srinivasakumar, Bartosz Stefaniak, Rahul Heinrich Steiger, Shangkun Solar, Jiaxiang Tang, Shitao Tang, Yangyang Tang, Yue Tang, Tolou Tavakkoli, Kayley Ting, Krzysztof Tomala, Wei-Cheng Tseng, Jibin Varghese, Sergei Vasilev, Thomas Volk, Raju Wagwani, Roger Waleffe, Andrew Z. Wang, Boxiang Wang, Haoxiang Wang, Qiao Wang, Shihao Wang, Shijie Wang, Ting-Chun Wang, Yan Wang, Yu Wang, David Wehr, Fangyin Wei, Xinshuo Weng, Jay Zhangjie Wu, Kedi Wu, Hongchi Xia, Summer season Xiao, Tianjun Xiao, Kevin Xie, Daguang Xu, Jiashu Xu, Mengyao Xu, Ruqing Xu, Xingqian Xu, Yao Xu, Dinghao Yang, Dong Yang, Hans Yang, Xiaodong Yang, Xuning Yang, Yichu Yang, Yurong You, Zhiding Yu, Hao Yuan, Simon Yuen, Xiaohui Zeng, Pengcuo Zeren, Cindy Zha, Haotian Zhang, Jenny Zhang, Jing Zhang, Liangkai Zhang, Paris Zhang, Shun Zhang, Xuanmeng Zhang, Zhizheng Zhang, Ann Zhao, Yilin Zhao, Yuliya Zhautouskaya, Charles Zhou, Fengzhe Zhou, Shilin Zhu, Yuke Zhu, Dima Zhylko, and Artur Zolkowski.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Bodily AI programs should perceive the actual world earlier than they will act inside it. Robots, autonomous autos, and sensible areas want to grasp what\u2019s taking place of their world, predict what\u2019s prone to occur subsequent, and generate actions for particular environments, embodiments, and duties. NVIDIA Cosmos 3 is a frontier basis mannequin for bodily [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":658,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Hammer-Robot-500x282.gif","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[954,955,952,293,81,953,208,135],"class_list":["post-656","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-action","tag-cosmos","tag-develop","tag-models","tag-nvidia","tag-physical","tag-reasoning","tag-world"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3 - Future News 24<\/title>\n<meta name=\"description\" content=\"Physical AI systems must understand the real world before they can act within it. Robots, autonomous vehicles, and smart spaces need to understand what&rsquo;s&#8230;\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3 - Future News 24\" \/>\n<meta property=\"og:description\" content=\"Physical AI systems must understand the real world before they can act within it. Robots, autonomous vehicles, and smart spaces need to understand what&rsquo;s&#8230;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-01T04:43:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-07T10:59:32+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Hammer-Robot-500x282.gif\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Hammer-Robot-500x282.gif\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3\",\"datePublished\":\"2026-06-01T04:43:00+00:00\",\"dateModified\":\"2026-06-07T10:59:32+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/\"},\"wordCount\":2167,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/Hammer-Robot-500x282.gif\",\"keywords\":[\"Action\",\"Cosmos\",\"Develop\",\"Models\",\"NVIDIA\",\"Physical\",\"Reasoning\",\"World\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/\",\"name\":\"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3 - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/Hammer-Robot-500x282.gif\",\"datePublished\":\"2026-06-01T04:43:00+00:00\",\"dateModified\":\"2026-06-07T10:59:32+00:00\",\"description\":\"Physical AI systems must understand the real world before they can act within it. Robots, autonomous vehicles, and smart spaces need to understand what&rsquo;s&#8230;\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/Hammer-Robot-500x282.gif\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/Hammer-Robot-500x282.gif\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/01\\\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3 - Future News 24","description":"Physical AI systems must understand the real world before they can act within it. Robots, autonomous vehicles, and smart spaces need to understand what&rsquo;s&#8230;","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/","og_locale":"en_US","og_type":"article","og_title":"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3 - Future News 24","og_description":"Physical AI systems must understand the real world before they can act within it. Robots, autonomous vehicles, and smart spaces need to understand what&rsquo;s&#8230;","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/","og_site_name":"Future News 24","article_published_time":"2026-06-01T04:43:00+00:00","article_modified_time":"2026-06-07T10:59:32+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Hammer-Robot-500x282.gif","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Hammer-Robot-500x282.gif","twitter_misc":{"Written by":"Future News 24","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3","datePublished":"2026-06-01T04:43:00+00:00","dateModified":"2026-06-07T10:59:32+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/"},"wordCount":2167,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Hammer-Robot-500x282.gif","keywords":["Action","Cosmos","Develop","Models","NVIDIA","Physical","Reasoning","World"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/","name":"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3 - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Hammer-Robot-500x282.gif","datePublished":"2026-06-01T04:43:00+00:00","dateModified":"2026-06-07T10:59:32+00:00","description":"Physical AI systems must understand the real world before they can act within it. Robots, autonomous vehicles, and smart spaces need to understand what&rsquo;s&#8230;","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Hammer-Robot-500x282.gif","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/05\/Hammer-Robot-500x282.gif"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/01\/develop-physical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Develop Bodily AI Reasoning, World, and Motion Fashions with NVIDIA Cosmos 3"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/656","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=656"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/656\/revisions"}],"predecessor-version":[{"id":657,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/656\/revisions\/657"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/658"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=656"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=656"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=656"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}