The paper introduces MISCO, an evolutionary framework that uses deep generative models to design voxel-based soft robots (VSRs). MISCO combines an estimation-of-distribution algorithm with a variational autoencoder that includes multi-task learning, position awareness, and inter-voxel signaling to improve representation and sampling efficiency. The authors provide theoretical guarantees of asymptotic convergence to globally optimal designs and demonstrate through simulations that MISCO effectively navigates large design spaces, producing high-performing VSRs for various tasks while balancing efficiency and diversity.
By Junru Song, Huan Xiao, Yang Yang, Guozhen Li, Wei Peng, Xiaoya Zhang, Tingsong Jiang, Weien Zhou, Ying Wen, Feifei Wang, Wen Yao
arXiv:2608.23100v1 Announce Type: cross
Abstract: Robot co-design via bi-level optimization couples within-lifetime controller learning for fitness evaluation with cross-generational morphological ev...
By Junru Song, Yang Yang, Yaqing Xu, Ying Wen, Wei Peng, Guozhen Li, Wei'en Zhou, Wen Yao
arXiv:2506. 08630v3 Announce Type: replace Abstract: A universal controller for any robot morphology would greatly improve computational and data efficiency.
By Laurens Engwegen, Max Weltevrede, Caroline Horsch, Daan Brinks, Wendelin B\"ohmer
arXiv:2609.25627v1 Announce Type: cross
Abstract: General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate prec...
By Haoran Wen, Wenfu Wang, Kunsong Shi, Jingke Wang, Wancheng Feng, Yiren Zhang, Yueran Zhao, Xuancheng Zhang, Nanfei Ye, Xingru Chen, Zhaohong Sun, Chengmin Yang, Zikang Yu, Penghao Bi, Jia Shi, Yu Liu, Kun Zhan, Yan Xie
arXiv:2606. 12109v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable zero-shot generalization in robotic manipulation, yet the vast majority of pre-trained pipelines remain strictly confined to low-DoF parallel grippers.
By Chuanke Pang, Junyi Huang, Zhijun Zhao, Yaobing Wang, Kun Xu, Xilun Ding
arXiv:2604. 21391v2 Announce Type: replace-cross Abstract: Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action.
By Yiming Zhong, Yaoyu He, Zemin Yang, Pengfei Tian, Yifan Huang, Qingqiu Huang, Xinge Zhu, Yuexin Ma
arXiv:2607. 05377v1 Announce Type: cross Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations.
By Jiaqi Peng, Xiqian Yu, Delin Feng, Yuqiang Yang, Wenzhe Cai, Jing Xiong, Ganlin Yang, Jinliang Zheng, Jiafei Cao, Xueyuan Wei, Jiangmiao Pang, Yuan Shen, Tai Wang
The paper investigates the LeWorldModel (LeWM) and its Sketched Isotropic Gaussian Regularizer (SIGReg), showing that the original Raw LeWM objective biases variance toward temporally persistent components, which suppresses residual variance and hampers robot state decodability. By applying SIGReg specifically to temporally centered residuals, the authors decouple persistent and residual variance allocation, improving representation quality. On the LIBERO benchmark, this adjustment boosts downstream policy success on the Goal suite by 1.66× and raises overall success rates from 63.6% to 83.8%, outperforming Diffusion Policy and pretrained OpenVLA without external pretraining.
By Chang Liu, Fei Suo, Yanzhou Jin, Zeyu Ping, Yusuke Iwasawa, Yutaka Matsuo, Yaonan Zhu
arXiv:2608. 05970v1 Announce Type: cross Abstract: Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks.
By Changyuan Wang, Chubin Zhang, Zhenyu Wu, Runhao Li, Angyuan Ma, Ke Chao, Yinan Liang, Xiuwei Xu, Ziwei Wang, Yansong Tang, Jiwen Lu
RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to evaluate Vision‑Language‑Action models on fine‑grained spatial reasoning and long‑horizon procedural planning. It contains 10 task categories, 56 base tasks, and 280 variants across five difficulty levels, with 527K trajectories collected from multiple embodiments and scenes. The benchmark introduces diagnostic metrics beyond binary success, revealing that current VLA models struggle with complex spatial relations, precise execution, and memory‑intensive planning.
By Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
arXiv:2605. 13548v3 Announce Type: replace-cross Abstract: Existing robotic foundation models, while powerful, are predicated on an implicit assumption of temporal homogeneity: treating all actions as equally informative during optimization.
By Daojie Peng, Fulong Ma, Jiahang Cao, Qiang Zhang, Xupeng Xie, Jian Guo, Ping Luo, Andrew F. Luo, Boyu Zhou, Jun Ma
arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.
By Yifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao, Yucheng Hu, Qiyu Wu, Yuxiao Li, Zibin Dong, Fei Ni, Yan Zheng, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye Hao