Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,858 stories · RSS feed

arXiv Computer Vision
Sep 16

Exploring 2D backbone effects for indoor semantic occupancy prediction

The paper investigates how different 2D image backbones affect indoor semantic occupancy prediction in RGB‑D pipelines. By keeping the projection, depth branch, and occupancy head constant and swapping only the image encoder, the authors find that stronger backbones such as DINOv2 and BLIP2 significantly raise mIoU scores compared to the default ResNet‑50. These results show that the choice of image backbone is a major determinant of 3D occupancy accuracy, outweighing many specialized 3D modules.

By Shizhang Fanga, Wanling Yea, Qi Zheng
arXiv Computer Vision
Sep 16

Racing in Volume with Flow Ensembles

The paper introduces FastFlowGS, a streaming 4D Gaussian Splatting method that reconstructs fast-moving subjects from a sparse set of external cameras, and Monaco4D, a photorealistic Unreal Engine 5 benchmark featuring Formula 1 sequences with dense ground truth. FastFlowGS combines sparse matches, semi-dense tracks, and dense optical flow using a Kalman-style temporal update, achieving significant performance gains over existing baselines on both CMU-Panoptic and Monaco4D datasets. The benchmark provides varied illumination and viewpoints from trackside, onboard, and drone cameras, enabling evaluation of high-speed outdoor reconstruction.

By Saswat Subhajyoti Mallick, Riu Cherdchusakulchai, Marc Ruiz Olle, Albert Mosella-Montoro, Jose Ribeiro-Gomes, Francisco Vicente Carrasco, Fernando De la Torre
arXiv Machine Learning
Sep 16

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.

By Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin
arXiv AI
Sep 16

Det-LIME: Detector-Aware, Multi-Instance Local Interpretable Model-Agnostic Explanations for Automated Marine Mammal Detection

Det‑LIME is a detector‑aware, multi‑instance adaptation of LIME designed to explain black‑box object detectors used in marine mammal research. It generates instance‑specific, box‑aligned explanations by weighting detections, applying a proximity kernel, and using IoU‑based matching to track instances across perturbations. Evaluated on aerial drone imagery of harbor seals and a seabird case study, Det‑LIME outperformed vanilla LIME, Stabilized LIME, Deterministic LIME, and gradient‑based methods in Attribution Ratio and Max Saliency Hit Rate, offering higher‑resolution, instance‑aware explanations that aid debugging, data augmentation, and modeling improvements.

By Jiayi Zhou, David W. Johnston, Brinnae Bent
arXiv Machine Learning
Sep 16

The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformer

The paper re‑examines the impact of removing the encoder from Action Chunking Transformers (ACT), a model used for robot manipulation learning. Contrary to the original claim that encoder removal drops success rates from 35% to 2%, the authors find no such dramatic effect in their re‑runs, though minor variations remain uncertain. They attribute discrepancies to training length and checkpoint selection, and note that the encoder’s latent variable offers little reconstruction benefit on the tested benchmark, while its removal speeds up training.

By Bo Kang
arXiv Machine Learning
Sep 16

TARC: Time-Adaptive Robotic Control

TARC (Time‑Adaptive Robotic Control) is a reinforcement‑learning framework that lets a policy predict both a control action and how long it should be applied, thereby learning temporally extended actions. By optimizing task performance under constraints on the number of control switches, TARC can adapt its control rate online, using high‑frequency feedback only when necessary. Experiments on a high‑speed RC car, a Unitree Go1 quadruped, and a vision‑language action model show that TARC matches the performance of high‑frequency discrete‑time controllers while operating at less than half their control frequency.

By Arnav Sukhija, Lenart Treven, Jin Cheng, Florian D\"orfler, Stelian Coros, Andreas Krause
arXiv Computer Vision
Sep 16

Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation

Mem-World introduces a memory‑augmented action‑conditioned world model for robot manipulation, featuring W‑VMem—a 4D wrist‑view‑centered surfel‑indexed memory that anchors historical observations to evolving surface elements. By explicitly modeling when and where scene elements are observed, the system retrieves geometry‑aware history frames during generation, providing informative, non‑redundant context for future action predictions. Experiments demonstrate that Mem‑World produces persistent rollouts, improves policy evaluation reliability (14.5 % higher Pearson correlation with real‑world performance), and boosts long‑horizon task success rates from 58 % to 72 % using synthetic data generation.

By Zirui Zheng, Jiaqian Yu, Xiongfeng Peng, jun shi, Mingyi Li, Chao Zhang, Weiming Li, Dong Wang, Huchuan Lu, Xu Jia
arXiv Computer Vision
Sep 16

SlotDiT: Object-Centric Representations for Diffusion Transformers

SlotDiT introduces a text-guided Diffusion Transformer that operates in a slot-based latent space, decomposing scenes into object-centric slots and autoregressively denoising future slot trajectories to predict scene dynamics. The model is conditioned on a reference image and a language instruction, enabling it to generate video content that reflects both visual context and textual guidance. Experiments comparing slot-based representations to VAE-based and semantics-aligned alternatives show that SlotDiT achieves competitive video generation quality while improving task-completion rates across four robotic datasets and offering a more computationally efficient latent representation.

By Gjergj Plepi, Sven Behnke
arXiv Computer Vision
Sep 16

World-Action Models for Robot Learning and Control: A Survey

The survey "World-Action Models for Robot Learning and Control" reviews recent advances in coupling future world prediction with executable action generation for robots in open environments. It clarifies the scope of World-Action Models (WAMs) relative to conventional world models, model-based RL, and Vision‑Language‑Action policies, and organizes existing methods through a unified taxonomy covering representations, transition modeling, action interfaces, architectures, training pipelines, data modalities, and scaling strategies. The paper also surveys applications in manipulation, navigation, and autonomous driving, summarizes datasets, benchmarks, and metrics, and discusses key challenges such as action alignment, spatial consistency, long‑horizon memory, and efficient inference.

By Zuxing Lu, Hongjia Zhai, Guanzhi Wang, Huajian Zeng, Jiaqi Yang, Jingyu Liu, Lei Cheng, Yuantai Zhang, Yuheng Qiu, Zezhou Cheng, Ivan Laptev, Danfei Xu, Benjamin Riviere, Giuseppe Loianno, Eric Xing, Xingxing Zuo