Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,943 stories · RSS feed

arXiv Computer Vision
Sep 15

FFVO: A Feedforward Pose Decoder for Long-Horizon Visual Odometry

arXiv:2609.13733v1 Announce Type: new Abstract: Stable and reliable 4D spatial understanding is fundamental for autonomous driving systems. While feedforward reconstruction networks can estimate came...

By Meng-Li Shih, Shih-Yang Su, Yuliang Zou, Hao Xiang, Haidong Zhu, Vincent Casser, Brian Curless, Dmitry Kalenichenko, Mingxing Tan, Dragomir Anguelov
arXiv Computer Vision
Sep 14

UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction

arXiv:2609.12898v1 Announce Type: new Abstract: Fine-grained robotic manipulation depends on understanding parts, not only whole objects. Existing 3D foundation models tend to be either generalized b...

By Xinqiang Yu, Zekun qi, Jiawei He, Wenyao Zhang, Xuchuan Chen, Guaocai Yao, Li Yi, Zhaoxiang Zhang, He Wang
arXiv Machine Learning
Sep 14

Decoding Mixture Perception through Computational Modeling of Component Interactions

The paper introduces a bio‑inspired deep learning framework that models olfactory perception of complex chemical mixtures. It constructs neural response curves for molecule‑receptor interactions, fuses attention‑weighted multi‑receptor and concentration‑dependent multi‑molecule data, and transfers knowledge from molecular associations to improve mixture recognition. The model achieves 92.2% accuracy and offers a generalizable computational pathway from chemical blending to perceptual formation.

By Fei Wang, Xiaoya Xie, Junfei Liu, Huihao Wang, Yixiao Wang, Yintao Wang, Yi Li, Hao Dong, Xing Chen
arXiv Computer Vision
Sep 14

MoPA: Coordinated Mobile Manipulation via Subsystem-Specific Perception Alignment

arXiv:2609.12081v1 Announce Type: cross Abstract: Mobile manipulation requires perceptual evidence at different spatial scales for base motion and arm control, while the two action modalities remain...

By Guangyu Chen, Qiwei Liang, Shaolong Zhu, Tianxing Chen, Zikuan Xiao, Yifan Xie, Lingfeng Zhang, Ping Luo, Renjing Xu, Wenbo Ding
arXiv Computer Vision
Sep 14

EgoFun3D: Modeling Interactive Objects from Egocentric Videos using Function Templates

EgoFun3D introduces a coordinated task, dataset, and benchmark for creating simulation-ready interactive 3D objects from egocentric videos. The approach captures general cross-part functional mappings via function templates, enabling precise evaluation and direct compilation into executable code. A 4‑stage pipeline—2D part segmentation, reconstruction, articulation estimation, and function template inference—is proposed, and a dataset of 517 videos with detailed annotations is released for comprehensive benchmarking.

By Weikun Peng, Denys Iliash, Manolis Savva
arXiv Machine Learning
Sep 14

VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search

VertexCBF is a framework that learns neural control barrier functions (CBFs) by approximating the stationary Hamilton–Jacobi value function with a neural network trained through physics‑informed and sparsely supervised learning. It exploits control‑affine dynamics and a convex polytope control set to generate supervision points via GPU‑parallel vertex‑restricted tree search, ensuring the learned CBF never exceeds the specified constraint function. The method was evaluated on 15 systems, outperforming baselines by recovering larger safe sets, and demonstrated on a mobile robot that safely avoids pedestrians using a neural CBF trained with this approach.

By Bojan Deraji\'c, Sebastian Bernhard, Wolfgang H\"onig
arXiv Machine Learning
Sep 14

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Dynin‑Robotics introduces an omnimodal masked‑diffusion backbone, Dynin‑Omni, that jointly represents language, visual observations, goals, and actions as discrete tokens. By conditioning on different spans, the same model learns action prediction, next‑observation prediction, goal‑state prediction, and trajectory‑to‑instruction reconstruction, enabling test‑time scaling through goal prediction and action‑candidate evaluation. The system, pretrained on 1.33 million trajectories from 48 Open X‑Embodiment datasets, achieves competitive performance on LIBERO, zero‑shot LIBERO‑Plus, and a 78.4 % success rate on a Franka Research 3 robot, while a block‑parallel implementation speeds up action decoding by up to 29.2×.

By Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do
arXiv Machine Learning
Sep 14

DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction

DenseTRF is a self‑supervised framework that adapts texture‑aware representations for dense prediction in surgical computer vision. It uses slot attention to learn invariant visual structures and then conditions dense prediction on these representations, merging models to adapt to target distributions without supervision. Experiments on multiple surgical procedures show that DenseTRF improves cross‑distribution generalization compared to state‑of‑the‑art segmentation models and test‑distribution adaptation methods.

By Guiqiu Liao, Matja\v{z} Jogan, Daniel A. Hashimoto
arXiv Machine Learning
Sep 14

Scenario-Independent Criticality Assessment and Prediction for Vulnerable Road Users in Autonomous Driving

The paper introduces a new criticality metric specifically designed for vulnerable road users (VRUs) and a scenario‑independent prediction framework that applies to all traffic participants. The VRU‑centric metric improves pedestrian criticality classification by up to 50 %, while the prediction framework surpasses state‑of‑the‑art metrics by 275 %, achieving an F1‑score of 0.96 on the DeepAccident dataset. These advances enable more accurate, scenario‑agnostic safety assessments for autonomous driving systems.

By J\"org Gamerdinger, Victor Schwarzenberger, Philipp Schmid, Sven Teufel, Oliver Bringmann
arXiv Machine Learning
Sep 14

Efficient Vision-Language-Action Management and Serving for Robot Factories

Robion is a new serving and management system designed to run Vision‑Language‑Action (VLA) models on multi‑GPU edge servers for robot factories. It splits the VLM and ADiT stages within a single GPU, shares streams across multiple models, and prioritizes requests by remaining SLO time, enabling high robot load while meeting strict latency requirements. In experiments, Robion achieves 6.7× higher robot load than vLLM‑Omni and 1.5× higher than a monolithic pipeline, and can serve 64 robots on a 4‑GPU server with 98% SLO attainment.

By Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri, Christina Giannoula
arXiv Machine Learning
Sep 14

The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile Devices

The paper investigates the environmental impact of running large language models (LLMs) on mobile devices. It evaluates 18 different LLM configurations on two smartphones and a server, measuring energy per token, latency, accuracy, and battery-cycle consumption. Findings reveal that on-device inference is about three times less energy‑efficient than batched server inference, that energy consumption varies non‑monotonically with quantization bit‑width, and that most models are not on the Pareto front of accuracy and energy efficiency. The study concludes that local AI is not inherently more sustainable than cloud inference, with the majority of environmental impact stemming from device embodied carbon.

By \'Edouard Gu\'egain, Tristan Coignion
arXiv Machine Learning
Sep 14

Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

The paper introduces CLAW, a method that uses a hypernetwork to generate low‑rank adapters for world models during test time, enabling efficient adaptation to new environments with only a few episodes of interaction. By jointly pretraining the hypernetwork and base model on simulated adaptations, CLAW balances computational efficiency and expressivity, outperforming both in‑context learning and gradient‑based adaptation in locomotion and manipulation tasks. The approach also mitigates overfitting in data‑scarce regimes and demonstrates that the benefit stems from expressive adapters rather than context conditioning.

By Fernando Palafox, David Fridovich-Keil
arXiv Machine Learning
Sep 14

Curriculum-Based Adversarial Heterogeneous Agent Reinforcement Learning for Autonomous Quad-Copter Landing in Maritime Settings

The paper presents a curriculum‑based adversarial heterogeneous agent reinforcement learning (HARL‑AC) approach for autonomous quad‑copter landing on a ship deck in maritime settings. Using Heterogeneous‑Agent Proximal Policy Optimization (HAPPO) in NVIDIA Isaac Lab, the authors train a cooperative control policy that outperforms domain‑randomized baselines, achieving up to 97.5% success on in‑distribution sea states and higher median success and lower crash rates on out‑of‑distribution sea states. The adversarially trained policy also exhibits more cautious behavior, slightly increasing timeouts but improving safety in severe, unseen conditions.

By Allan Minh-Tam Nguyen, Sree Showrya Kotala, Stefan Banioi-Crijman, Kurt Driessens, Rico M\"ockel