arXiv:2606. 16190v1 Announce Type: cross Abstract: Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or privacy constraints.
By Zhihan Zhang, Alexander Le Metzger, Jiuyang Lyu, Chun-Cheng Chang, Jiayi Shao, Yujia Liu, Emmanuel Azuh Mensah, Edward Wang, Kurtis Heimerl, Gregory D. Abowd, Shwetak Patel, Natasha Jaques, Vikram Iyer
The paper presents a hardware‑accelerated instance segmentation framework tailored for resource‑constrained lunar robotics, addressing low‑light perception, limited compute, and radiation‑induced hardware faults. It introduces Activation Variance Informative Sampling (AVIS), a label‑free calibration method that selects samples based on activation variance, and deploys a YOLO‑based model on a Deep Learning Processor Unit with architectural tweaks to reduce CPU fallback and ensure bounded latency. A software‑level criticality analysis estimates fault exposure, guiding mitigation that reduces global criticality by 31.7%, while AVIS with bias correction recovers 69.8% of quantization‑induced accuracy loss at 309 ms latency and 5.7 W power consumption.
By Siddhant Shete, Hilmi Dogu K\"uc\"uker, Udo Frese, Frank Kirchner
arXiv:2505. 01458v2 Announce Type: replace-cross Abstract: Navigation and manipulation are core capabilities in Embodied AI, but training agents to perform them directly in the real world is costly, time-consuming, and unsafe.
By Lik Hang Kenny Wong, Xueyang Kang, Kaixin Bai, Jianwei Zhang
We’re introducing an efficient, on-device robotics model with general-purpose dexterity and fast task adaptation.
arXiv:2607. 00710v1 Announce Type: cross Abstract: Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones.
By Richard Schwarzkopf, Jonas Merkert, Frank Bieder, Annika B\"atz, Alexander Blumberg, Carlos Fernandez, Felix Hauser, Fabian Immel, Christian Kinzig, Hendrik K\"onigshof, Fabian Konstantinidis, Martin Lauer, Willi Poh, Nils Rack, Kevin R\"osch, Yinzhe Shen, Marlon Steiner, Gleb Stepanov, Dominik Strutz, \"Omer \c{S}ahin Ta\c{s}, Julian Truetsch, Kaiwen Wang, Royden Wagner, Jan-Hendrik Pauls, Christoph Stiller
arXiv:2606. 00162v1 Announce Type: cross Abstract: Robotic systems generate large volumes of multimodal sensor data, but converting ROS bag recordings into machine learning datasets is often handled by ad hoc sequential scripts, creating engineering overhead and slow iteration cycles.
By Leon Pohl, Lukas Beer, George Sebastian, Mirko Maehlisch
StarVLA-α is a streamlined Vision‑Language‑Action (VLA) model that reduces architectural and pipeline complexity to facilitate systematic study of VLA design choices. By employing a strong VLM backbone and minimal design, it achieves competitive performance across multiple benchmarks (LIBERO, SimplerEnv, RoboTwin, RoboCasa) and outperforms the baseline π₀.₅ by 20% on the RoboChallenge benchmark. The authors plan to release the code to support future VLA research.
By Jinhui Ye, Ning Gao, Senqiao Yang, Jinliang Zheng, Zixuan Wang, Yuxin Chen, Pengguang Chen, Yilun Chen, Shu Liu, Jiaya Jia
FluxVLA Engine is an open, configuration‑driven platform that unifies the fragmented components of embodied policy development—datasets, visual‑language and world models, action heads, learning methods, distributed training, simulation evaluation, inference, and robot interfaces—into a reproducible data‑to‑deployment workflow. It adds features such as compositional dual‑arm simulation, scalable automatic data generation, human‑in‑the‑loop rollout and correction, Real‑Time Chunking for fast inference, and lightweight remote GPU serving, thereby linking offline learning, simulation validation, online correction, and real‑robot execution under shared, auditable contracts. The engine aims to eliminate engineering bottlenecks that currently separate promising embodied‑learning algorithms from reliable, reproducible deployment.
By Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong, Qirui Hu, Yiming Zhang, Weipeng Deng, Bowen Shen, Minzhao Zhu, Yiming Mao, Yan Yang, Chenguang Cui, Hongyuan Chen, Xu Huang, Zheyi Zhao, Pinxi Shen, Bozhen He, Zhen Fu, Yifan Wang, Zexin Zhang, Ang Gao, Haoyu Chen, Chengqi Shi, Hua Chen
The paper introduces Self‑Adaptive VLA, a post‑training method that lets Vision‑Language‑Action policies self‑adapt to deployment‑time hardware shifts by using rollouts as context. It creates shift‑conditioned expert demonstrations, compresses visual, proprioceptive, and action data into a latent context token, and modulates the policy via adaptive layer normalization. Experiments on four precision‑critical manipulation tasks show the method recovers over 80 % of the base policy’s performance under actuation bias and encoder offsets, and improves robustness on new workstations.
By Hongxin Zhang, Chunru Lin, Tsun-Hsuan Wang, Zhenjia Xu, Chuang Gan
arXiv:2608. 19891v1 Announce Type: new Abstract: How to efficiently finetune robot policies to learn new tasks on the fly?
By Bhavya Sukhija, Oliver Groth, Mohit Shridhar, Tim Hertweck, Michael Bloesch, Markus Wulfmeier, Abbas Abdolmaleki, Martin Riedmiller
arXiv:2603. 11811v2 Announce Type: replace-cross Abstract: The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by the prohibitive cost and scalability limits of human-in-the-loop collection paradigms.
By Yongzhong Wang, Keyu Zhu, Yong Zhong, Liqiong Wang, Jinyu Yang, Feng Zheng