Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,943 stories · RSS feed

arXiv AI
Sep 17

A Comprehensive Review of Generative Physical Artificial Intelligence

The paper surveys Generative Physical Artificial Intelligence (GPAI), a field where large foundation models are integrated with physical robots. It introduces a taxonomy of five approaches—Robot Foundation Models, Vision‑Language Action models, Large Behavior Models, Diffusion Policy Models, and World Foundation Models—and discusses how they complement each other across domains such as autonomous vehicles, industrial automation, healthcare robotics, and humanoid systems. The review highlights performance gains, data‑efficient learning, sim‑to‑real transfer, edge‑compatible architectures, and safety frameworks as key research directions.

By Satyam Gaba, Krutiksinh Rana, Siva Sai, Vinay Chamola, Dusit Niyato
arXiv AI
Sep 17

Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation

Mem2Ego introduces a vision‑language model for embodied navigation that combines global memory with egocentric visual inputs. By adaptively retrieving task‑relevant cues from a global memory module and aligning them with local perception, the framework improves spatial reasoning and decision‑making over long horizons. The method outperforms prior state‑of‑the‑art approaches on the HSSD and HM3D benchmarks and shows strong performance on a real robot.

By Lingfeng Zhang, Yuecheng Liu, Zhanguang Zhang, Matin Aghaei, Yixin Xiao, Yaochen Hu, Mohammad Ali Alomrani, David Gamaliel Arcos Bravo, Hongjian Gu, Zhiyuan Li, Yangzheng Wu, Zhanpeng Zhang, Raika Karimi, Atia Hamidizadeh, Guowei Huang, Haoping Xu, Tongtong Cao, Weichao Qiu, Xingyue Quan, Jianye Hao, Yuzheng Zhuang, Yingxue Zhang
arXiv Computer Vision
Sep 17

Multi-View Mixture-of-Experts with Vision-Language Reranking for Cross-View Object Geo-Localization

The paper introduces MVLGeo, a unified framework for cross-view object geo-localization that combines multiple viewpoints into a single model. It employs Vision‑Language Reranking to use contextual text from the query view, a multi‑view Mixture‑of‑Experts architecture to share knowledge and reduce redundancy, and an adaptive elliptical prior for positional encoding. Experiments on CVOGL benchmarks show that MVLGeo achieves state‑of‑the‑art performance and robustness to input degradation.

By Xuyu Fan, Qi Ming, Zhu Han, Liuqian Wang, Siyuan Cao, Xiaohan Zhang, Xudong Zhao, Mingjing Zhao, Yuhan Zhang
arXiv AI
Sep 17

Imitation Learning for Autonomous Driving in CARLA

The paper presents a compact multimodal policy trained via behavioral cloning to drive autonomously in the CARLA simulator. Using five‑frame histories of RGB images, LiDAR, telemetry, and lane waypoints, the 1.36‑million‑parameter model predicts throttle, brake, and steering at 20 Hz. Trained on 236,882 windows (≈3.3 hours of driving) from 448 captures, the policy drives for hours on both training and unseen routes without collisions, demonstrating qualitative transfer and recovery from large trajectory deviations.

By Jordy Kieto
arXiv AI
Sep 17

Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents

The paper introduces Market Signal Injection (MSI), an attack that alters how market data is formatted or described—without changing its numerical values—to influence large language model (LLM) pricing agents. Experiments on nine open‑weight and three proprietary models in simulated duopoly and triopoly markets show that sentiment‑based formatting changes cause significant shifts in firm behavior, profits, and consumer surplus. The study also demonstrates that model susceptibility varies across families, that larger models are not always more robust, and that techniques such as input canonicalization and decision boundary anchoring can partially mitigate these attacks.

By Dohun Lee, Hyunwoo Park
arXiv Machine Learning
Sep 17

The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformers

The paper re‑examines the impact of removing the encoder from Action Chunking Transformers (ACT), a model used for robot manipulation learning. Contrary to the original claim that encoder removal drops success rates from 35% to 2%, the authors find no such dramatic decline in their re‑runs, though minor variations remain uncertain. They attribute the discrepancy to factors like training length and checkpoint selection, and note that the encoder’s latent variable offers little reconstruction benefit on the tested benchmark, while its removal speeds up training.

By Bo Kang
arXiv Machine Learning
Sep 17

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

By Andrew P. Berg, Qian Zhang, Mia Y. Wang
arXiv Machine Learning
Sep 17

Learning from Distributed Eyes: Leveraging Collaborative Perception for Automated Model Adaptation

The paper introduces LDE, a framework that uses collaborative perception (CP) to generate high‑quality pseudo‑labels for unsupervised model adaptation in autonomous driving. It tackles communication limits, field‑of‑view mismatches, and label unreliability through selective feature sharing, FoV filtering, and curriculum learning. Experiments on 3D object detection show LDE surpasses pre‑trained models and existing adaptation methods.

By Yanan Ma, Yihang Tao, Zhengru Fang, Zihan Fang, Yiqin Deng, Xianhao Chen, Yuguang Fang
arXiv Computer Vision
Sep 17

LiteViLNet: Lightweight Vision-LiDAR Fusion Network for Efficient Road Segmentation

LiteViLNet is a lightweight RGB‑geometry fusion network for road segmentation that uses a MobileNetV3 RGB encoder and a tiny depth‑wise‑separable geometry encoder. Its multi‑scale fusion module enhances modality‑specific features, performs cross‑modal interaction, and applies adaptive gating, while a depth‑wise large‑kernel bridge expands contextual support with minimal overhead. The U‑Net‑style decoder is trained with deep supervision, achieving state‑of‑the‑art performance on KITTI and ORFD benchmarks and running at up to 68.73 FPS on a Jetson Orin NX with TensorRT FP16.

By Daojie Peng, Bingtao Wang, Fulong Ma, Liang Zhang, Jun Ma
arXiv Computer Vision
Sep 17

HAP: A Hand-Driven Active Perception Framework for Egocentric Head Motion Prediction

The paper introduces HAP, a Hand-Driven Active Perception framework that predicts future six‑degree‑of‑freedom head motion in egocentric settings by conditioning on observed hand motion and inferred target context. HAP constructs a Predictive Target‑Centric Amodal Occlusion Graph to model current and potential occlusions among candidate objects, fuses this with hand and head motion history, and blends the learned trajectory with a constant‑velocity prior. Experiments on a public dataset and a newly released Bottle RGB‑D dataset demonstrate that HAP outperforms baseline methods in head‑motion prediction, highlighting the importance of hand‑driven intention and dynamic occlusion reasoning.

By Yunji Feng, Junyi Ma, Guanzhong Sun, Chenyang Xu, Hesheng Wang
arXiv Computer Vision
Sep 17

PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image

PhysVGGT is a feed‑forward model that predicts dense maps of friction coefficient, Shore hardness, Young's modulus, and density, along with object‑level mass, from a single RGB image in one forward pass. It treats physical property estimation as a dense per‑pixel prediction problem, using a visual geometry transformer to extract geometry‑aware tokens and separate dense and global prediction branches. A scalable pseudo‑label generation pipeline enables large‑scale weakly supervised training, and the model achieves state‑of‑the‑art performance on the ABO‑500 dataset while running 27× faster than previous methods.

By Sneha Paul, Guile Wu, Bingbing Liu, Dongfeng Bai
arXiv AI
Sep 17

AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution

AeroWeaver is a new embodied‑agent harness that integrates large language model (LLM) decision making with the executable skills of individual UAVs, enabling distributed, adaptive swarm execution. It connects semantic mission decisions to governed skills, organizes role‑conditioned local agents for coordination, and refines skill selection online using role‑indexed state‑action‑reward experience. Experiments demonstrate that AeroWeaver maintains valid skill execution without a central joint‑action generator and supports reward‑guided, training‑free adaptive learning from accumulated execution experience.

By Jiabin Lou, Yirong Yang, Haopeng Wang, Xuxin Lv, Xinyu Liu, Diyuan Hou, Xuehong Liu, Rongye Shi, Wenjun Wu
arXiv AI
Sep 17

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

The paper introduces rMuscle, a real‑time Vision‑Language‑Action inference framework that mimics human muscle memory to accelerate robotic decision making. By exploiting repeated task similarity, rMuscle uses a dual‑phase cache: a Context Cache reuses visual‑token outputs and an Action Cache reuses neuron activation patterns, reducing computation and weight accesses. Experiments on RTX 4090 and Jetson Thor show 1.29–1.42× speedups on LIBERO, RoboTwin, and physical manipulation tasks while preserving success rates on real robots.

By Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu
arXiv AI
Sep 17

Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents

SynAgent is a framework that uses large language model agents to run autonomous experiments while building an explicit, revisable understanding of the synthesis process. Unlike traditional black‑box optimizers, SynAgent generates analysis skills on the fly and reasons multimodally over data such as X‑ray diffraction patterns and electron micrographs. In an 18‑experiment campaign on LiCoO₂ thin‑film deposition, SynAgent produced highly crystalline films and uncovered a sharp temperature threshold and optimal growth window (650–690 °C) for crystallization.

By Izumi Takahara, Kazunori Nishio, Akira Aiba, Shigeru Kobayashi, Takao Nakajima, Taro Hitosugi, Teruyasu Mizoguchi
arXiv AI
Sep 17

PACT-WAM: Predicting Actions and Visual Foresight with Compact Temporal Encoding for Robot Manipulation

PACT‑WAM is a world‑action model that simultaneously predicts a 16‑step action trajectory and its corresponding visual forecast for robot manipulation. It uses a hierarchical history encoder that compresses past observations into fewer tokens, reducing processing cost by 75% compared to dense encoding. The model’s shared flow module updates action and visual states jointly, and a TiTok‑VAE decoder reconstructs multi‑view future images, which are then used by a vision‑language component (Proposal Review) to improve execution‑prefix selection and proposal rejection, boosting success rates on several benchmarks.

By Yushan Liu, Jingjing Fan, Shoujie Li, Yifan Xie, Xiao-Ping Zhang, Wenbo Ding
arXiv Machine Learning
Sep 17

Changepoint-Aware World Models: Detecting Dynamics Shifts and Recovering by Forgetting Stale Replay in Model-Based RL

Changepoint-Aware World Models (CAWM) is a DreamerV3 agent that detects abrupt dynamics shifts in a robot’s environment using an online CUSUM test on internal prediction error. Upon detection, CAWM selectively forgets stale replay data while preserving the learned representation, enabling rapid recovery from shifts such as doubled gravity or halved actuator gain. Experiments on simulated locomotion show CAWM recovers faster than passive retraining and outperforms a baseline that respawns a fresh dynamics model, achieving significant return gains in the first 30k post‑shift frames.

By Everest Yang
arXiv Computer Vision
Sep 17

Energy-Regularized Imitation Learning for Force- and Work-Aware Robotic Manipulation

This paper introduces an energy-aware approach to robotic manipulation by defining a joint-space mechanical-work proxy based on joint torque and angular displacement. A differentiable energy predictor is trained to estimate this work from robot states and actions, enabling it to serve as a regularizer that fine‑tunes a pretrained manipulation policy. Applied to RVT‑2 on RLBench, the method reduces average mechanical work from 208.8 J to 204.4 J (a 2.1 % drop) while slightly improving task success from 86.2 % to 86.9 % across 12 manipulation tasks.

By Toshiki Otani, Hiromu Taketsugu, Norimichi Ukita