Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv Machine Learning
Sep 25

Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

The paper introduces TRACE, an amortized temporal gradient‑inversion attack that reconstructs private observation‑action trajectories from per‑step policy gradients in embodied reinforcement‑learning agents. TRACE exploits cross‑time correlation between gradients and exact action recovery from policy‑head gradients, achieving high reconstruction quality (18.8 dB PSNR) and near‑perfect action recovery with minimal computation. The study demonstrates TRACE’s effectiveness across various neural architectures and input modalities, and suggests that protecting temporal gradient streams may require sequence‑aware privacy mechanisms.

By Sudip Bhujel, Shanghao Shi, Ruiquan Huang, Ning Zhang, Yang Xiao
arXiv AI
Sep 25

Towards An LLM-Driven Unified Conversion Framework for BT and FSM in Autonomous Intelligent Systems

The paper introduces an LLM-driven framework that automatically converts between finite state machines (FSM) and behavior trees (BT) for autonomous intelligent systems. It addresses key challenges such as preserving behavioral completeness and preventing model complexity explosion by designing a loop execution BT structure and employing depth compression strategies with LLM prompts. Experiments in various autonomous decision-making scenarios show that the framework achieves accurate, scalable, and maintainable bidirectional conversion, benefiting consumer-grade applications like service robots, game agents, and smart home devices.

By Zhang Qi, Yang Shuo, Zhu Zhengqiu, Zhou Peng, Jiao Peng
arXiv AI
Sep 25

PUBG Ally: A Conversational Embodied Agent as an AI Teammate

PUBG Ally is an embodied, voice‑enabled AI teammate for PUBG: BATTLEGROUNDS that can perceive the game world, interpret player speech, and autonomously decide actions while keeping speech synchronized with gameplay. It combines a language‑model agent that uses a controlled interface to gather game information and a faster control layer for movement, combat, and recovery. The system was trained on nearly 39,000 real‑player sessions and evaluated through player feedback and preference comparisons, with live deployment requiring low‑latency on‑device execution and safety safeguards.

By Beomsoo Kim, Byeongju Kim, Dohyun Kim, Dongwon Kim, Eunchong Kim, Hongmin Kim, Hyeojung Im, Hyeonbin Hwang, Hyeonghwan Kim, Hyoseok Seol, Insub Im, Irene Chen, Jaeseung Jeon, Jimin Hong, Kiyoon Yoo, Minkyoung Park, Seohyeon Jung, Seungjun Chung, Sue Hyun Park, Sungwoo Kim, Youngin Cho, Yujeong Son, Kangwook Lee, Hyunseung Kim
arXiv AI
Sep 25

Design and Evaluation of LLM Chaining-Based Task Planning for General Purpose Service Robots

The paper introduces an LLM chaining architecture for General Purpose Service Robots that splits instruction classification and action generation into two stages, cutting prompt length by about 45% and boosting planning consistency. Evaluation on 100 synthetic GPSR commands across three language models shows consistent improvements over single-prompt methods, with up to +37 percentage points gain on local models. Real‑robot trials on the Toyota HSR confirm that while planning success improves, execution-layer failures remain the main obstacle to full task completion.

By Lucas Da Mota Bruno, Jiahao Sim, Yoshinobu Hagiwara
arXiv AI
Sep 25

EIB-Net: Entropy-Guided Information Bottleneck for Generalizable AI-Generated Image Detection

EIB-Net is an Entropy‑Guided Information Bottleneck Network designed to detect AI‑generated images across diverse generative models. It introduces an Image Entropy metric to automatically select the most informative, low‑entropy patch and applies a Variational Information Bottleneck to learn compact, generalizable features. Experiments on DIFF, DiffusionForensics, and GenImage benchmarks show state‑of‑the‑art performance, achieving 85.7% accuracy with only 2% of training data and maintaining robust cross‑generator generalization.

By Zhida Zhang, Xinlei Ma, Jie Cao
arXiv AI
Sep 25

DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models

DAWN (Denoising and Alignment in World models for Noise-robustness) is a perception framework that builds noise robustness directly into a world model for vision-based legged locomotion. It achieves this by feeding noisy depth to the encoder while reconstructing clean depth, and by using contrastive learning to align latent states of noisy and clean depth. The method requires no manual filter tuning, incurs no extra inference cost, and enables zero‑shot quadruped parkour on a Unitree Go1, successfully traversing stairs, gaps, and steps from raw depth observations.

By Yohan Choi, Min-Jun Kim, Jin-Sung Kim, Yong-Jae Kim, Youn-Hee Han
arXiv AI
Sep 25

HarnessPAI: An Evolving Harness for Physical AI

HarnessPAI is a model‑ and embodiment‑agnostic framework that treats code as an executable, evolvable interface for Physical AI. It separates short‑term open‑loop program execution from long‑term closed‑loop evolution, using feedback to refine programs and distill reusable skills. Across various robots, HarnessPAI outperforms pure action models and code‑as‑policy baselines, achieving significant gains on tasks like LIBERO‑PRO and RoboCasa without retraining the underlying model.

By Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan, Yang Li, Qing Li, Shangding Gu, Huichi Zhou, Shuqing Shi, Fei Ni, Shuo Lu, Weicheng Meng, Kang Li, Jin Wu, Kang Zhao, Shangmin Guo, Gen Li, Yongqiang Tang, Zhizhong Zhang, Yuan Xie, Heng Qu
arXiv AI
Sep 25

Operator Packages, Proposer Strength, and Construction-Family Plateaus in Office-Scale Verified Search

The paper reports on a large‑scale verified search experiment using a 30B language model on a laptop, evaluating three operator packages—schematic notebooks, named obstacles, and behavioural repulsion—in a factorial design across nine construction problems. Results show that the full composition of operators closes the seed‑to‑record gap more effectively than any single component, increases construction‑hash diversity, and that memory plus repulsion consistently avoids collapse. A frontier proposer achieves similar gains in far fewer samples, but the search ultimately stalls near a plateau where the reference family is adopted and optimized only when provided as code.

By Roberto I. Ono Filho
arXiv AI
Sep 25

When Search Becomes Memory: Accelerating Robot Design Discovery with Self-Evolving Skills

The paper introduces Auto‑Robotist, a self‑evolving large language model (LLM) agent that transforms evolutionary robot design search traces into an explicit natural‑language skill library. Each skill records a structural archetype, evidence‑grounded rules, and supporting designs, enabling the agent to retrieve and condition LLM edits during search while still using a genetic algorithm for exploration. Experiments on seven EvoGym tasks show that Auto‑Robotist outperforms standard genetic algorithms, especially when transferring learned skills to larger design spaces.

By Yunfei Wang, Xiaohao Xu, Yang Li, Xiaonan Huang
arXiv Machine Learning
Sep 25

Error- and Prediction-Driven Motor Learning in the Cortico-Cerebellar Loop

The paper presents a cerebellum-inspired control framework that addresses robust control under delayed sensory feedback. By combining multiplexed predictive representations with internal feedback, the model jointly encodes kinematic variables and task-relevant error signals, enabling accurate online correction despite delays. Incorporating feedback within the cerebellar loop also accelerates adaptation, reducing learning time by an order of magnitude.

By Ana Carolina Filipe, Rui Ponte Costa, Cl\'audia Soares
arXiv Computation and Language
Sep 25

Free the Language Model From the Vision Encoder: Semantic Serialization as a Perception Interface for Small Language Models

The paper introduces a perception interface that separates vision from language in vision‑language models. A frozen perception stack detects objects, a deterministic semantic serializer converts the perceived state into text, and a standard text‑only large language model (LLM) answers questions. Experiments on a campus‑robot benchmark show that this serialized interface outperforms a zero‑shot VLM of the same language‑model size, especially as the language model shrinks, and that the advantage persists under paraphrase and different supervision regimes.

By Cong Xu, Ravi Sankar
arXiv Computer Vision
Sep 25

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

M3GD introduces a multimodal representation that fuses pre‑trained 2D image and 3D LiDAR foundation models for robotic novel view synthesis, avoiding the need for a separate cross‑modal translator. By projecting LiDAR onto the image latent grid and injecting the resulting geometry‑aware packets via a lightweight residual adapter, the method enhances both RGB and depth synthesis on the GrandTour dataset compared to an image‑only baseline. Ablation studies confirm that pixel‑aligned LiDAR content drives the performance gains, and real‑world deployment on a ground robot demonstrates a tunable quality–cost trade‑off.

By Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, Giuseppe Loianno
arXiv Computer Vision
Sep 25

RotVLA: Rotational Latent Action for Vision-Language-Action Model

RotVLA introduces a Vision‑Language‑Action framework that replaces discrete latent action encoding with a continuous rotational latent action representation on the group SO(n). This design provides continuity, compositionality, and structured geometry that better capture real‑world action dynamics, and a triplet frame learning scheme enforces meaningful temporal dynamics while preventing degeneration. Trained with 1.7 B parameters on large cross‑embodiment datasets, RotVLA achieves state‑of‑the‑art performance on LIBERO and RoboTwin2.0 benchmarks and shows strong real‑world manipulation results.

By Qiwei Li, Xicheng Gong, Xinghang Li, Peiyan Li, Quanyun Zhou, Hangjun Ye, Jiahuan Zhou, Yadong Mu
arXiv AI
Sep 25

Novelty Adaptation Through Hybrid Large Language Model (LLM)-Symbolic Planning and LLM-guided Reinforcement Learning

The paper introduces a neuro‑symbolic architecture that combines symbolic planning, reinforcement learning, and a large language model (LLM) to address novelties in dynamic open‑world environments. The LLM is used to identify missing operators, generate symbolic plans, and write reward functions, enabling the reinforcement learning agent to learn control policies for newly identified operators. The proposed method outperforms state‑of‑the‑art approaches in both operator discovery and learning within continuous robotic domains.

By Hong Lu, Pierrick Lorang, Timothy R. Duggan, Jivko Sinapov, Matthias Scheutz
arXiv Computer Vision
Sep 25

Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures

Ego-Exo4D is a large-scale dataset that offers synchronized egocentric and multi-view exocentric video for applications such as skill learning, procedural activity understanding, and embodied AI. The original dataset only includes sparse 3D human pose annotations, making dense motion reconstruction challenging. To address this, the authors introduce Ego-Exo4D-HM, a new dataset containing 4D human motion reconstructions for the Ego-Exo4D captures, along with a reconstruction pipeline and accompanying code and documentation.

By Abhiram Maddukuri, Georgios Pavlakos
Hugging Face Trending Papers
Sep 24

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

M3GD is a novel approach for robotic novel view synthesis that fuses camera images and LiDAR point clouds without requiring a separate cross‑modal translator. By projecting LiDAR data onto the image latent grid and injecting it via a lightweight residual adapter, M3GD enhances both RGB and depth generation on the GrandTour dataset compared to image‑only baselines. Experiments on a ground robot confirm that the method can be deployed in real‑world scenarios with a tunable quality‑cost trade‑off.

Hugging Face Trending Papers
Sep 24

S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

S2Planner is a trajectory planner for autonomous driving that fuses data from three front-facing cameras, ego‑motion history, and the current driving command. It uses a fine‑tuned DINOv3 backbone with a Spatial Tuning Adapter to generate multi‑scale image features, which are then refined by a coarse‑to‑fine decoder employing trajectory self‑attention and camera‑projected cross‑attention. The key contribution lies in integrating ego‑conditioned trajectory initialization with iterative, geometry‑guided sampling of multi‑scale image features, rather than introducing a new visual backbone or attention operator.

Hugging Face Trending Papers
Sep 24

HarnessPAI: An Evolving Harness for Physical AI

HarnessPAI introduces a model‑ and embodiment‑agnostic harness framework for Physical AI that treats code as an executable, evolvable interface organizing action primitives. The framework operates on two timescales: within a rollout it executes an open‑loop program, and across rollouts it evolves a closed‑loop program using feedback to refine the program and distill reusable skills. Across diverse robotic platforms, HarnessPAI outperforms pure action models and code‑as‑policy baselines, achieving significant gains on tasks such as LIBERO‑PRO and RoboCasa, and enabling efficient expert‑data collection for further fine‑tuning.

arXiv Machine Learning
Sep 24

Less Language, More Latents: Annotation-Efficient VLAs for Driving

The paper introduces Latent Action Driving Annotations (LADA), a three‑stage pipeline that converts large amounts of unlabelled observation‑trajectory data into a language‑conditioned driving model. First, a latent action model with a vector‑quantised bottleneck learns a compact codebook of vehicle intents. Then, a small set of language‑annotated examples trains a vision‑language translator to map observations and instructions into this codebook, and finally a VLA is trained on observation‑latent‑action pairs across the full corpus. Using less than 5% of language annotations, LADA attains a Driving Score of 87.98 and a Success Rate of 70.46% on Bench2Drive, matching or surpassing fully supervised baselines.

By Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania