arXiv AI

RoboLDA: A Probabilistic Generative Model for Uncovering Embodied Hierarchical Structures in Voxel-based Soft Robots

RoboLDA is a Bayesian probabilistic model that learns a four‑level hierarchy—task, robot, organ, voxel—from existing high‑performing voxel‑based soft robot designs. By training with variational inference, it uncovers consistent, intuitive hierarchical patterns and can generate new robot morphologies that outperform evolutionary algorithms by an average of 106.4% without further optimization. The inferred organ structures also improve modular control policies, demonstrating the model’s utility for zero‑shot design and motion control.

arXiv AI
Sep 25

Generative Evolutionary Design of Voxel-Based Soft Robots with Provable Optimality

The paper introduces MISCO, an evolutionary framework that uses deep generative models to design voxel-based soft robots (VSRs). MISCO combines an estimation-of-distribution algorithm with a variational autoencoder that includes multi-task learning, position awareness, and inter-voxel signaling to improve representation and sampling efficiency. The authors provide theoretical guarantees of asymptotic convergence to globally optimal designs and demonstrate through simulations that MISCO effectively navigates large design spaces, producing high-performing VSRs for various tasks while balancing efficiency and diversity.

By Junru Song, Huan Xiao, Yang Yang, Guozhen Li, Wei Peng, Xiaoya Zhang, Tingsong Jiang, Weien Zhou, Ying Wen, Feifei Wang, Wen Yao
arXiv Computer Vision
Sep 23

MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

arXiv:2609.25627v1 Announce Type: cross Abstract: General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate prec...

By Haoran Wen, Wenfu Wang, Kunsong Shi, Jingke Wang, Wancheng Feng, Yiren Zhang, Yueran Zhao, Xuancheng Zhang, Nanfei Ye, Xingru Chen, Zhaohong Sun, Chengmin Yang, Zikang Yu, Penghao Bi, Jia Shi, Yu Liu, Kun Zhan, Yan Xie
arXiv AI
Jun 11

Bridging the Morphology Gap: Adapting VLA Models to Dexterous Manipulation via Intent-Conditioned Fine-Tuning

arXiv:2606. 12109v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable zero-shot generalization in robotic manipulation, yet the vast majority of pre-trained pipelines remain strictly confined to low-DoF parallel grippers.

By Chuanke Pang, Junyi Huang, Zhijun Zhao, Yaobing Wang, Kun Xu, Xilun Ding
arXiv AI
Jun 16

From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges

arXiv:2604. 21391v2 Announce Type: replace-cross Abstract: Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action.

By Yiming Zhong, Yaoyu He, Zemin Yang, Pengfei Tian, Yifan Huang, Qingqiu Huang, Xinge Zhu, Yuexin Ma
arXiv AI
Jul 7

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

arXiv:2607. 05377v1 Announce Type: cross Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations.

By Jiaqi Peng, Xiqian Yu, Delin Feng, Yuqiang Yang, Wenzhe Cai, Jing Xiong, Ganlin Yang, Jinliang Zheng, Jiafei Cao, Xueyuan Wei, Jiangmiao Pang, Yuan Shen, Tai Wang
arXiv Machine Learning
Aug 27

Temporally Centered SIGReg Improves LeWorldModel Representations for Robot Policy Learning

The paper investigates the LeWorldModel (LeWM) and its Sketched Isotropic Gaussian Regularizer (SIGReg), showing that the original Raw LeWM objective biases variance toward temporally persistent components, which suppresses residual variance and hampers robot state decodability. By applying SIGReg specifically to temporally centered residuals, the authors decouple persistent and residual variance allocation, improving representation quality. On the LIBERO benchmark, this adjustment boosts downstream policy success on the Goal suite by 1.66× and raises overall success rates from 63.6% to 83.8%, outperforming Diffusion Policy and pretrained OpenVLA without external pretraining.

By Chang Liu, Fei Suo, Yanzhou Jin, Zeyu Ping, Yusuke Iwasawa, Yutaka Matsuo, Yaonan Zhu
arXiv AI
Aug 7

SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

arXiv:2608. 05970v1 Announce Type: cross Abstract: Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks.

By Changyuan Wang, Chubin Zhang, Zhenyu Wu, Runhao Li, Angyuan Ma, Ke Chao, Yinan Liang, Xiuwei Xu, Ziwei Wang, Yansong Tang, Jiwen Lu
arXiv AI
Sep 7

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to evaluate Vision‑Language‑Action models on fine‑grained spatial reasoning and long‑horizon procedural planning. It contains 10 task categories, 56 base tasks, and 280 variants across five difficulty levels, with 527K trajectories collected from multiple embodiments and scenes. The benchmark introduces diagnostic metrics beyond binary success, revealing that current VLA models struggle with complex spatial relations, precise execution, and memory‑intensive planning.

By Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
arXiv AI
Jun 11

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.

By Yifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao, Yucheng Hu, Qiyu Wu, Yuxiao Li, Zibin Dong, Fei Ni, Yan Zheng, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye Hao