Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,943 stories · RSS feed

arXiv Machine Learning
Sep 16

LCAP: Population-Informed Latent Chip Adaptation from Few Output Probes for Photonic Neural Networks

The paper introduces LCAP, a method for adapting photonic neural networks to real hardware by learning a shared correction from a population of chips and then personalizing each chip using only 32 fixed output probes. LCAP decomposes adaptation into a transferable population correction and a probe‑inferred latent personalization, allowing feed‑forward calibration without device‑specific optimization. Experiments on a simulated three‑layer 64‑mode MZI network show accuracy improvements from 80.4% to 93.4% and significant gains on unseen chips.

By Tianyu Gao, Guantian Zheng
arXiv Computer Vision
Sep 16

Occupancy Network-Guided Autonomous Robotic Partial Nephrectomy

arXiv:2609.16186v1 Announce Type: cross Abstract: Autonomous soft-tissue cancer surgery has been limited to interventions on organ surfaces, because current systems cannot perceive and adapt to anato...

By Ethan Kilmer, Pit Henrich, Jiawei Ge, Paul M. Scheikl, Laura Connolly, Soum D. Lokeshwar, Joseph Chen, Justin D. Opfermann, Kaitlyn Kumar, Lauren Shepard, Ahmed Ghazi, Nirmish Singla, Richard J. Cha, Kevin Cleary, Franziska Mathis-Ullrich, Axel Krieger
arXiv Computer Vision
Sep 16

DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding

DGSG-Mind introduces a hybrid instance-aware 3D Gaussian dynamic scene graph system that integrates open‑vocabulary semantic information into dynamic 3D scene representations. By coupling a probabilistic voxel grid with explicit 3D Gaussians, it achieves robust cross‑modal instance fusion, incremental semantic mapping, and dynamic change handling through Gaussian‑based relocalization and masked refinement. The system builds a hierarchical scene graph and a 3D Gaussian Mind for multimodal reasoning, achieving state‑of‑the‑art zero‑shot 3D visual grounding and strong performance in open‑vocabulary semantic segmentation and scene reconstruction, and is demonstrated on real‑world robots.

By Luzhou Ge, Xiangyu Zhu, Jinyan Liu, Xuesong Li
arXiv Computer Vision
Sep 16

World-Action Models for Robot Learning and Control: A Survey

The survey "World-Action Models for Robot Learning and Control" reviews recent advances in coupling future world prediction with executable action generation for robots in open environments. It clarifies the scope of World-Action Models (WAMs) relative to conventional world models, model-based RL, and Vision‑Language‑Action policies, and organizes existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. The paper also surveys applications in manipulation, navigation, and autonomous driving, summarizes datasets, benchmarks, and metrics, and discusses key challenges such as action alignment, spatial consistency, long‑horizon memory, and efficient inference.

By Zuxing Lu, Hongjia Zhai, Guanzhi Wang, Huajian Zeng, Jiaqi Yang, Jingyu Liu, Lei Cheng, Yuantai Zhang, Yuheng Qiu, Zezhou Cheng, Ivan Laptev, Danfei Xu, Benjamin Riviere, Giuseppe Loianno, Eric Xing, Xingxing Zuo
arXiv Computer Vision
Sep 16

SlotDiT: Object-Centric Representations for Diffusion Transformers

SlotDiT introduces a text-guided Diffusion Transformer that operates in a slot-based latent space, decomposing scenes into object-centric slots and autoregressively denoising future slot trajectories to predict scene dynamics. The model is conditioned on a reference image and a language instruction, enabling it to generate video content that reflects both visual context and textual guidance. Experiments comparing slot-based representations to VAE-based and semantics-aligned alternatives show that SlotDiT achieves competitive video generation quality while improving task-completion rates across four robotic datasets and offering a more computationally efficient latent representation.

By Gjergj Plepi, Sven Behnke
arXiv Computer Vision
Sep 16

Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation

Mem-World introduces a memory‑augmented action‑conditioned world model for robot manipulation, featuring W‑VMem—a 4D wrist‑view‑centered surfel‑indexed memory that anchors historical observations to evolving surface elements. By explicitly modeling when and where scene elements are observed, the system retrieves geometry‑aware history frames during generation, providing informative, non‑redundant context for future action predictions. Experiments demonstrate that Mem‑World produces persistent rollouts, improves policy evaluation reliability (14.5 % higher Pearson correlation with real‑world performance), and boosts long‑horizon task success rates from 58 % to 72 % using synthetic data generation.

By Zirui Zheng, Jiaqian Yu, Xiongfeng Peng, jun shi, Mingyi Li, Chao Zhang, Weiming Li, Dong Wang, Huchuan Lu, Xu Jia
arXiv AI
Sep 16

World Models for Embodied Intelligence: From Plausible to Controllable to Actionable

arXiv:2609.16697v1 Announce Type: cross Abstract: World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interv...

By Nanjie Yao, Hao Wang, Chong Cheng, Zhikang Chen, Wenzhe Li, Jiafei Lyu, Li Shen, Peilin Zhao, Zongqing Lu, Gao Huang, Steven Hoi, Dacheng Tao, Deheng Ye
arXiv Machine Learning
Sep 16

TARC: Time-Adaptive Robotic Control

TARC (Time‑Adaptive Robotic Control) is a reinforcement‑learning framework that lets a policy predict both a control action and how long it should be applied, thereby learning temporally extended actions. By optimizing task performance under constraints on the number of control switches, TARC can adapt its control rate online, using high‑frequency feedback only when necessary. Experiments on a high‑speed RC car, a Unitree Go1 quadruped, and a vision‑language action model show that TARC matches the performance of high‑frequency discrete‑time controllers while operating at less than half their control frequency.

By Arnav Sukhija, Lenart Treven, Jin Cheng, Florian D\"orfler, Stelian Coros, Andreas Krause
arXiv Computer Vision
Sep 16

Differentiable Mesh State Estimation via Factor Graph Inference for Deformable Object Reconstruction

arXiv:2609.16686v1 Announce Type: cross Abstract: Estimating deformable object states remains a fundamental challenge in robotics and simulation. We propose a novel factor graph-based framework for p...

By Lidia Al-Zogbi, Fangjie Li, Samuel Tobin, James Ferguson, Nithesh Kumar, Alejandro Chara, Kuan-I Chung, Mingxing Rao, Ayberk Acar, Susheela Sharma Stern, Robert Webster, Daniel Moyer, Alan Kuntz, Caleb Rucker, Tucker Hermans, Jie Ying Wu
arXiv Computer Vision
Sep 16

GeoLAM: Learning Geometry-Grounded Latent Actions from Unlabeled Human Videos

GeoLAM is a framework that learns geometry‑grounded latent actions from unlabeled human videos. It uses future‑frame reconstruction with a frozen geometric feature hierarchy and motion supervision from a 4D geometry teacher to capture 3D displacement, image‑plane motion, and surface‑orientation changes. After pretraining, the representation serves as transition targets for a world‑action model trained on robot demonstrations, enabling denoised latent actions and executable action chunks without requiring hand‑pose annotations or future‑video generation during deployment.

By Yifan Xie, Hekun Tian, Jinkun Liu, YuAn Wang, Qiao Sun, Wenbo Ding
arXiv Computer Vision
Sep 16

MEgoVista: Multi-view Ego-aware Motion Estimation for Metric 4D Hands and Head in the Wild

MEgoVista is an offline pipeline that converts a single unprepared egocentric video into metric two‑hand and head motion within a gravity‑aligned world frame. It uniquely reconstructs motion in environments beyond studio volumes, uses calibrated stereo for absolute scale, and evaluates its outputs against independent optical capture to audit accuracy. The system thus expands the settings where high‑fidelity hand‑motion labels can be generated from natural, head‑worn recordings.

By Jiangong Xiao (Northwestern Polytechnical University), Zhihao Zhang (Xi'an Jiaotong University), Yifei Dong (Maniformer), Chao Ma (Maniformer), Zhouyi Jin (Maniformer), Zhiwen Hou (Maniformer), Li Liu (Maniformer), Weihuang Chen (Xi'an Jiaotong University), Hongbin Sun (Xi'an Jiaotong University), Maoqing Yao (Maniformer)
arXiv Computer Vision
Sep 16

EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset

EventEgoHands++ is a new framework for reconstructing 3D hand meshes from egocentric event-based cameras. It introduces a Hand Detector that provides instance-level bounding boxes and masks for left and right hands, and an Adaptive Attention module that uses these detections to model spatial relationships and interactions. The authors extend the synthetic N-HOT3D dataset and create EEH‑R, a large real-world event-based egocentric hand dataset with about 1 million annotated frames, and show that their method outperforms existing baselines on both synthetic and real data.

By Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa