To function successfully throughout various contexts, robots should not solely carry out manipulation duties precisely but in addition adapt how their actions unfold to the duty, object, and interplay setting. We ask whether or not this execution-level variation might be realized as a reusable behavioral issue shared throughout duties. We current MoMo, a two-stage imitation-learning framework consisting of a spatiotemporal motion tokenizer and a behavior-cloning transformer that takes activity and a steady motion-mode situation as inputs. Throughout six real-robot manipulation duties, various this situation produces regular, dynamic, and intermediate behaviors that human raters can distinguish and that differ in joint velocity, acceleration, and end-effector strategy pitch. On duties demonstrated in just one mode, MoMo transfers the unseen requested mode whereas largely preserving activity success. Collectively, these outcomes present proof of compositional generalization to unseen activity–mode mixtures and present that movement mode might be reused throughout duties to regulate how a manipulation talent is carried out.

