{"id":3305,"date":"2026-08-04T16:00:00","date_gmt":"2026-08-04T16:00:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/"},"modified":"2026-08-05T05:59:10","modified_gmt":"2026-08-05T05:59:10","slug":"beyond-vlas-how-world-action-models-reshape-robot-manipulation","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/","title":{"rendered":"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\">A central problem in robotics is constructing insurance policies that generalize past the demonstrations they\u2019re educated on. A coverage that succeeds in a coaching scene usually fails when object shapes, positions, or lighting change. Generalizing to those new situations requires the coverage to grasp the duties underlying physics, not simply mimic the demonstrations. This potential comes from the spine it\u2019s constructed on.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The usual solution to construct a language-conditioned robotic coverage is so as to add an motion module to a pretrained vision-language mannequin (VLM), producing a vision-language-action (VLA) mannequin. This strategy has carried generalist manipulation a good distance. However a VLM spine learns to explain the world, not predict the way it evolves. That lacking dynamics mannequin is precisely what a robotic wants when a activity depends upon anticipating how a scene will change. A rising line of analysis replaces the language spine with a video world mannequin, producing world motion mannequin (WAM). NVIDIA researcher Jim Fan explored this shift in his Robotics\u2019 Finish Sport discuss\u2014an concept later summarized as \u201cVLAs are useless, lengthy dwell World Motion Fashions.\u201d<\/p>\n<p class=\"wp-block-paragraph\">This submit explores how post-training can flip WAMs into specialised robotic coverage, how WAMs examine to VLAs, and why the open NVIDIA Cosmos 3 world mannequin offers a robust basis for constructing WAMs.<\/p>\n<h2 id=\"how_post-trained_policies_are_built_today\" class=\"wp-block-heading\">How post-trained insurance policies are constructed at the moment<\/h2>\n<p class=\"wp-block-paragraph\">Within the VLA paradigm, a pretrained VLM offers semantic understanding of scenes and language directions, whereas post-training learns to map that understanding to robotic actions. Fashionable generalist robotic insurance policies more and more construct on this strategy.<\/p>\n<p class=\"wp-block-paragraph\">Nevertheless, a VLM is optimized to provide textual content about pictures, to not mannequin how\u00a0 a scene will evolve. It doesn&#8217;t be taught what occurs to a mug when the gripper closes, how a towel folds, the place an object lands when launched. VLAs generalize nicely semantically however are much less efficient at bodily generalization to unseen behaviors and environments, as their backbones mannequin language and notion quite than world dynamics.\u00a0<\/p>\n<h2 id=\"what_a_wam_changes\" class=\"wp-block-heading\">What a WAM modifications<\/h2>\n<p class=\"wp-block-paragraph\">A WAM overcomes that dynamics-modeling restrict by constructing the coverage on a video world mannequin. As a result of the spine fashions how the world evolves, post-training doesn&#8217;t have to show dynamics from scratch, it specializes a mannequin that already has a physics prior. The NVIDIA analysis paper World Motion Fashions are Zero-shot Insurance policies exhibits that collectively predicting video and motion provides a coverage properties a VLA can&#8217;t simply purchase:<\/p>\n<p>It learns from numerous knowledge. A VLA learns to map directions to trajectories, and infrequently requires near-identical demonstrations of the identical activity. A WAM learns physics: how objects transfer when pushed, grasped, or dropped. Any interplay knowledge teaches it one thing. A diversified dataset that may be wasted on a VLA turns into a coaching sign, and knowledge assortment will get cheaper.<\/p>\n<p>It generalizes within the open world. Physics is extra common than semantics. The best way an object falls or slides doesn\u2019t change with the article, so a mannequin that has realized dynamics carries that data into scenes and motions it was by no means educated on.<\/p>\n<p>It adapts to new robots with few demonstrations. A mannequin that already understands bodily interplay wants far much less task-specific knowledge to specialize to a brand new arm or gripper.<\/p>\n<p class=\"wp-block-paragraph\">For a group constructing a coverage, these are the sensible wins: much less knowledge to succeed in a given functionality, higher habits exterior the coaching distribution, and a shorter path to a brand new embodiment. They&#8217;re properties of the pretraining spine, in order that they present up in each coverage post-trained from it.<\/p>\n<h2 id=\"cosmos_3_is_a_strong_wam_foundation\" class=\"wp-block-heading\">Cosmos 3 is a robust WAM basis<\/h2>\n<p class=\"wp-block-paragraph\">Cosmos 3 is an omni-model world basis mannequin constructed on a Combination-of-Transformers (MoT) structure. Multimodal enter flows by way of an autoregressive transformer for reasoning producing discrete tokens corresponding to textual content. This guides a diffusion transformer for steady modalities, together with picture, video, audio, and motion. Textual content is generated by next-token decoding; every thing else, together with actions, is synthesized by way of iterative denoising. A single mannequin spans these modalities whereas retaining the era mechanism finest suited to every. Cosmos 3 is available in three sizes: 4B NVIDIA Cosmos Edge, 16B NVIDIA Cosmos Nano, and 64B NVIDIA Cosmos 3 Tremendous.<\/p>\n<p class=\"wp-block-paragraph\">What makes Cosmos 3 a robust basis for post-training is the breadth of its physical-world knowledge. The dataset consists of roughly 767M pictures,348M movies of real-world dynamics, 8M motion samples spanning robotic manipulation, autonomous driving, digicam movement, and selfish movement.\u00a0<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a72d114e2573&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a72d114e2573\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"2048\" height=\"1078\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85.webp\" alt=\"Comparison of VLA and world action model architectures. A VLA passes video and text through a VLM reasoner, then combines its output with robot state in a diffusion action head to generate actions. A world action model jointly post-trains a reasoner and diffusion model on video, text, and state to predict both future video frames and robot actions.\" class=\"wp-image-120684\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85.webp 2048w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-179x94.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-300x158.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-768x404.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-625x329.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-1536x809.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-645x340.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-500x263.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-160x84.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-362x191.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-209x110.png 209w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-1024x539.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-960x505.png 960w\" sizes=\"(max-width: 2048px) 100vw, 2048px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"2048\" height=\"1078\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85.webp\" alt=\"Comparison of VLA and world action model architectures. A VLA passes video and text through a VLM reasoner, then combines its output with robot state in a diffusion action head to generate actions. A world action model jointly post-trains a reasoner and diffusion model on video, text, and state to predict both future video frames and robot actions.\" class=\"lazyload wp-image-120684\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85.webp 2048w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-179x94.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-300x158.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-768x404.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-625x329.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-1536x809.png 1536w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-645x340.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-500x263.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-160x84.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-362x191.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-209x110.png 209w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-1024x539.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-85-960x505.png 960w\" data-sizes=\"(max-width: 2048px) 100vw, 2048px\"\/><figcaption class=\"wp-element-caption\">Determine 1. A VLA generates robotic actions from semantic reasoning and robotic state, whereas a WAM collectively predicts actions and future world states<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"from_world_model_to_robot_policy_cosmos_3_policy_droid_models\" class=\"wp-block-heading\">From world mannequin to robotic coverage: Cosmos 3 Coverage DROID fashions<\/h2>\n<p class=\"wp-block-paragraph\">Cosmos 3 is the start line for specialization. Cosmos3-Nano-Coverage-DROID is a 16B-parameter coverage post-trained from Cosmos 3 Nano for the DROID platform, which is a Franka Panda arm with a Robotiq gripper.A 4B model, Cosmos3-Edge-Coverage-DROID, is post-trained the identical manner, which can be utilized for on-device deployment. Given a language instruction and multi-camera observations, it generates robotic motion trajectories. Three properties observe from its omni basis:<\/p>\n<p>It imagines whereas it acts. When the mannequin outputs actions, it could possibly additionally output a video: what the robotic\u2019s cameras will see if these actions are executed. The motion and the anticipated consequence come from the identical mannequin, on the similar time.<\/p>\n<p>It retains the complete omni structure. Submit-training removes nothing. The coverage checkpoint can nonetheless cause and generate video, not simply output joint positions.<\/p>\n<p>The prior is measurable. The Cosmos 3 technical report compares two DROID insurance policies educated with the identical recipe,\u00a0 knowledge, and compute. One began from the bottom checkpoint, whereas the opposite began from an omni checkpoint educated on multi-domain motion knowledge. The omni checkpoint raised RoboLab success from 28.1% to 36.8%. That is clear proof that the development comes from the structure, not simply from scale.<\/p>\n<div class=\"wp-block-image\">\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a72d114e33ce&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a72d114e33ce\" class=\"aligncenter size-full wp-lightbox-container\"><img decoding=\"async\" width=\"1382\" height=\"415\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86.webp\" alt=\"Robot gripper positioned above a table with a banana and red bowl. Beside the camera view, the model describes a plan to grasp the banana and place it in the bowl, while a seven-channel line graph shows the corresponding continuous action sequence, including the gripper closing.\" class=\"wp-image-120685\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86.webp 1382w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-179x54.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-300x90.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-768x231.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-625x188.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-645x194.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-500x150.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-160x48.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-362x109.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-366x110.png 366w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-1024x307.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-960x288.png 960w\" sizes=\"(max-width: 1382px) 100vw, 1382px\"\/><img loading=\"lazy\" decoding=\"async\" width=\"1382\" height=\"415\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86.webp\" alt=\"Robot gripper positioned above a table with a banana and red bowl. Beside the camera view, the model describes a plan to grasp the banana and place it in the bowl, while a seven-channel line graph shows the corresponding continuous action sequence, including the gripper closing.\" class=\"lazyload wp-image-120685\" srcset=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86.webp 1382w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-179x54.png 179w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-300x90.png 300w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-768x231.png 768w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-625x188.png 625w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-645x194.png 645w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-500x150.png 500w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-160x48.png 160w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-362x109.png 362w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-366x110.png 366w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-1024x307.png 1024w, https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/image-86-960x288.png 960w\" data-sizes=\"(max-width: 1382px) 100vw, 1382px\"\/><figcaption class=\"wp-element-caption\">Determine 2. A robotic observes a banana on the desk (left); the mannequin\u2019s predicted motion plan is proven as textual content\u00a0<\/figcaption><\/figure>\n<\/div>\n<h2 id=\"deployment_considerations\" class=\"wp-block-heading\">Deployment issues<\/h2>\n<p class=\"wp-block-paragraph\">A WAM carries the complete generative world mannequin, not simply an motion head. It\u2019s bigger than compact VLAs, however Cosmos 3 map to totally different deployment tiers quite than forcing one trade-off:<\/p>\n<p>Workstation serving (Nano, 16B ). Cosmos3-Nano-Coverage runs beside the robotic quite than on board. Actual-world DROID deployment serves it on a single NVIDIA RTX PRO 6000, with the robotic streaming observations over the community and receiving motion chunks again.<\/p>\n<p>On-device (Edge, 4B). Cosmos 3 Edge runs the identical coverage workload immediately on embedded {hardware}. It operates at robot-control decision (640\u00d7360 observations) and generates 32 actions per inference on NVIDIA Jetson Thor whereas attaining real-time management at 15 Hz. It\u2019s supported throughout NVIDIA edge computer systems together with RTX PRO GPUs, DGX, GeForce RTX GPUs, and Jetson, together with the brand new Jetson T2000 and T3000 modules.<\/p>\n<h2 id=\"why_build_robot_policies_with_cosmos_3\" class=\"wp-block-heading\">Why construct robotic insurance policies with Cosmos 3?<\/h2>\n<p class=\"wp-block-paragraph\">WAMs symbolize a shift from studying to behave to studying how the world evolves. Cosmos 3 makes this strategy sensible:<\/p>\n<p>Open basis, SOTA place to begin. The bottom mannequin, datasets, post-training recipe, educated weights, analysis instruments, and serving stack are all launched below a license that allows business use.<\/p>\n<p>Sooner adaptation. Sturdy bodily priors minimize the task-specific knowledge a brand new coverage wants. Convert your knowledge to the LeRobotDataset format the robot-learning ecosystem already data into, run the revealed recipe.<\/p>\n<p>One basis, many robots. Every new embodiment, corresponding to Franka, dual-arm setups, UR, WidowX, nonetheless requires its personal post-training run. However all of them begin from the identical pretrained basis quite than pretraining a world mannequin from scratch, lowering the demonstrations wanted for every.<\/p>\n<h2 id=\"get_started\" class=\"wp-block-heading\">Get began<\/h2>\n<p class=\"wp-block-paragraph\">The quickest solution to consider whether or not a WAM beats your present VLA is to post-train one from Cosmos 3 by yourself knowledge and examine.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/developer.nvidia.com\/blog\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>A central problem in robotics is constructing insurance policies that generalize past the demonstrations they\u2019re educated on. A coverage that succeeds in a coaching scene usually fails when object shapes, positions, or lighting change. Generalizing to those new situations requires the coverage to grasp the duties underlying physics, not simply mimic the demonstrations. This potential [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3307,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/Cosmos-World-Models-Robot.gif","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[3],"tags":[954,3566,293,3706,1206,1942,135],"class_list":["post-3305","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-platforms-apps","tag-action","tag-manipulation","tag-models","tag-reshape","tag-robot","tag-vlas","tag-world"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Past VLAs: How World Motion Fashions Reshape Robotic Manipulation - Future News 24<\/title>\n<meta name=\"description\" content=\"A central challenge in robotics is building policies that generalize beyond the demonstrations they&rsquo;re trained on. A policy that succeeds in a training scene&#8230;\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation - Future News 24\" \/>\n<meta property=\"og:description\" content=\"A central challenge in robotics is building policies that generalize beyond the demonstrations they&rsquo;re trained on. A policy that succeeds in a training scene&#8230;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-04T16:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-05T05:59:10+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/Cosmos-World-Models-Robot.gif\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/Cosmos-World-Models-Robot.gif\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"6 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation\",\"datePublished\":\"2026-08-04T16:00:00+00:00\",\"dateModified\":\"2026-08-05T05:59:10+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/\"},\"wordCount\":1304,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/Cosmos-World-Models-Robot.gif\",\"keywords\":[\"Action\",\"Manipulation\",\"Models\",\"Reshape\",\"Robot\",\"VLAs\",\"World\"],\"articleSection\":[\"AI Platforms &amp; Apps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/\",\"name\":\"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/Cosmos-World-Models-Robot.gif\",\"datePublished\":\"2026-08-04T16:00:00+00:00\",\"dateModified\":\"2026-08-05T05:59:10+00:00\",\"description\":\"A central challenge in robotics is building policies that generalize beyond the demonstrations they&rsquo;re trained on. A policy that succeeds in a training scene&#8230;\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/#primaryimage\",\"url\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/Cosmos-World-Models-Robot.gif\",\"contentUrl\":\"https:\\\/\\\/developer-blogs.nvidia.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/Cosmos-World-Models-Robot.gif\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/08\\\/04\\\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation - Future News 24","description":"A central challenge in robotics is building policies that generalize beyond the demonstrations they&rsquo;re trained on. A policy that succeeds in a training scene&#8230;","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/","og_locale":"en_US","og_type":"article","og_title":"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation - Future News 24","og_description":"A central challenge in robotics is building policies that generalize beyond the demonstrations they&rsquo;re trained on. A policy that succeeds in a training scene&#8230;","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/","og_site_name":"Future News 24","article_published_time":"2026-08-04T16:00:00+00:00","article_modified_time":"2026-08-05T05:59:10+00:00","og_image":[{"url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/Cosmos-World-Models-Robot.gif","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/Cosmos-World-Models-Robot.gif","twitter_misc":{"Written by":"Future News 24","Est. reading time":"6 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation","datePublished":"2026-08-04T16:00:00+00:00","dateModified":"2026-08-05T05:59:10+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/"},"wordCount":1304,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/Cosmos-World-Models-Robot.gif","keywords":["Action","Manipulation","Models","Reshape","Robot","VLAs","World"],"articleSection":["AI Platforms &amp; Apps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/","name":"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/#primaryimage"},"thumbnailUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/Cosmos-World-Models-Robot.gif","datePublished":"2026-08-04T16:00:00+00:00","dateModified":"2026-08-05T05:59:10+00:00","description":"A central challenge in robotics is building policies that generalize beyond the demonstrations they&rsquo;re trained on. A policy that succeeds in a training scene&#8230;","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/#primaryimage","url":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/Cosmos-World-Models-Robot.gif","contentUrl":"https:\/\/developer-blogs.nvidia.com\/wp-content\/uploads\/2026\/07\/Cosmos-World-Models-Robot.gif"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/08\/04\/beyond-vlas-how-world-action-models-reshape-robot-manipulation\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"Past VLAs: How World Motion Fashions Reshape Robotic Manipulation"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3305","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=3305"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3305\/revisions"}],"predecessor-version":[{"id":3306,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/3305\/revisions\/3306"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/3307"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=3305"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=3305"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=3305"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}