Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,943 stories · RSS feed

arXiv Computer Vision
Sep 14

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

AnchorVLN is an open‑vocabulary vision‑language navigation system that separates semantic proposals from geometric metrics. It uses a VLM to generate semantics while a geometry module supplies reliable metric quantities such as range and bearing, all within a Model Context Protocol server. The system achieves 64.4% on instruction following and improves object‑reference accuracy, reducing median center error from 3.37 m to 2.48 m.

By Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti
arXiv Computer Vision
Sep 14

UniMo: Unifying Human and Animal Motion Generation

UniMo introduces a unified point‑cloud based framework for generating 3D motion that works for both humans and animals, overcoming challenges posed by diverse skeletal topologies and limited animal datasets. It converts parametric skeletons into unparametric representations and uses dynamic sampling to focus on active joints. The authors also release UniML3D, a large motion‑language dataset with 145,907 sequences and 433,388 captions, and demonstrate state‑of‑the‑art performance on multiple benchmarks.

By Zeyu Zhang, Zhiyuan Zhang, Siheng Wang, Yiran Wang, Danning Li, Ian Reid, Richard Hartley
arXiv Machine Learning
Sep 14

VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search

VertexCBF is a framework that learns neural control barrier functions (CBFs) by approximating the stationary Hamilton–Jacobi value function with a neural network trained through physics‑informed and sparsely supervised learning. It exploits control‑affine dynamics and a convex polytope control set to generate supervision points via GPU‑parallel vertex‑restricted tree search, ensuring the learned CBF never exceeds the specified constraint function. The method was evaluated on 15 systems, outperforming baselines by recovering larger safe sets, and demonstrated on a mobile robot that safely avoids pedestrians using a neural CBF trained with this approach.

By Bojan Deraji\'c, Sebastian Bernhard, Wolfgang H\"onig
arXiv Machine Learning
Sep 14

Decoding Mixture Perception through Computational Modeling of Component Interactions

The paper introduces a bio‑inspired deep learning framework that models olfactory perception of complex chemical mixtures. It constructs neural response curves for molecule‑receptor interactions, fuses attention‑weighted multi‑receptor and concentration‑dependent multi‑molecule data, and transfers knowledge from molecular associations to improve mixture recognition. The model achieves 92.2% accuracy and offers a generalizable computational pathway from chemical blending to perceptual formation.

By Fei Wang, Xiaoya Xie, Junfei Liu, Huihao Wang, Yixiao Wang, Yintao Wang, Yi Li, Hao Dong, Xing Chen
arXiv Machine Learning
Sep 14

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Dynin‑Robotics introduces an omnimodal masked‑diffusion backbone, Dynin‑Omni, that jointly represents language, visual observations, goals, and actions as discrete tokens. By conditioning on different spans, the same model learns action prediction, next‑observation prediction, goal‑state prediction, and trajectory‑to‑instruction reconstruction, enabling test‑time scaling through goal prediction and action‑candidate evaluation. The system, pretrained on 1.33 million trajectories from 48 Open X‑Embodiment datasets, achieves competitive performance on LIBERO, zero‑shot LIBERO‑Plus, and a 78.4 % success rate on a Franka Research 3 robot, while a block‑parallel implementation speeds up action decoding by up to 29.2×.

By Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do
arXiv Machine Learning
Sep 14

Efficient Vision-Language-Action Management and Serving for Robot Factories

Robion is a new serving and management system designed to run Vision‑Language‑Action (VLA) models on multi‑GPU edge servers for robot factories. It splits the VLM and ADiT stages within a single GPU, shares streams across multiple models, and prioritizes requests by remaining SLO time, enabling high robot load while meeting strict latency requirements. In experiments, Robion achieves 6.7× higher robot load than vLLM‑Omni and 1.5× higher than a monolithic pipeline, and can serve 64 robots on a 4‑GPU server with 98% SLO attainment.

By Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri, Christina Giannoula
arXiv Computation and Language
Sep 14

Agent as Policy for Robotic Manipulation

arXiv:2609.12541v1 Announce Type: new Abstract: We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-spec...

By Mengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang
arXiv Computer Vision
Sep 14

RoMu4o: A Robotic Manipulation Unit For Orchard Operations Automating Proximal Hyperspectral Leaf Sensing

RoMu4o is a ground robot equipped with a 6‑DOF arm and a vision system that performs real‑time deep‑learning image processing and motion planning for proximal hyperspectral leaf sensing in orchards. The system uses robust perception and manipulation pipelines to identify leaf 3D structure, propose 6‑D poses, and generate collision‑free, constraint‑aware paths for precise leaf grasping and spectroscopy. In lab trials the robot achieved a 95 % success rate for 1‑LPB hyperspectral sampling, while field trials in a pistachio orchard reached 70 % success for autonomous leaf grasping and measurement. whyItMatters":"The system demonstrates a viable robotic solution to automate leaf‑level hyperspectral sensing, addressing labor shortages and enabling precise crop health monitoring in precision agriculture."

By Mehrad Mortazavi, David J. Cappelleri, Reza Ehsani
arXiv Machine Learning
Sep 14

Curriculum-Based Adversarial Heterogeneous Agent Reinforcement Learning for Autonomous Quad-Copter Landing in Maritime Settings

The paper presents a curriculum‑based adversarial heterogeneous agent reinforcement learning (HARL‑AC) approach for autonomous quad‑copter landing on a ship deck in maritime settings. Using Heterogeneous‑Agent Proximal Policy Optimization (HAPPO) in NVIDIA Isaac Lab, the authors train a cooperative control policy that outperforms domain‑randomized baselines, achieving up to 97.5% success on in‑distribution sea states and higher median success and lower crash rates on out‑of‑distribution sea states. The adversarially trained policy also exhibits more cautious behavior, slightly increasing timeouts but improving safety in severe, unseen conditions.

By Allan Minh-Tam Nguyen, Sree Showrya Kotala, Stefan Banioi-Crijman, Kurt Driessens, Rico M\"ockel
arXiv Machine Learning
Sep 14

Scenario-Independent Criticality Assessment and Prediction for Vulnerable Road Users in Autonomous Driving

The paper introduces a new criticality metric specifically designed for vulnerable road users (VRUs) and a scenario‑independent prediction framework that applies to all traffic participants. The VRU‑centric metric improves pedestrian criticality classification by up to 50 %, while the prediction framework surpasses state‑of‑the‑art metrics by 275 %, achieving an F1‑score of 0.96 on the DeepAccident dataset. These advances enable more accurate, scenario‑agnostic safety assessments for autonomous driving systems.

By J\"org Gamerdinger, Victor Schwarzenberger, Philipp Schmid, Sven Teufel, Oliver Bringmann
arXiv Machine Learning
Sep 14

The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile Devices

The paper investigates the environmental impact of running large language models (LLMs) on mobile devices. It evaluates 18 different LLM configurations on two smartphones and a server, measuring energy per token, latency, accuracy, and battery-cycle consumption. Findings reveal that on-device inference is about three times less energy‑efficient than batched server inference, that energy consumption varies non‑monotonically with quantization bit‑width, and that most models are not on the Pareto front of accuracy and energy efficiency. The study concludes that local AI is not inherently more sustainable than cloud inference, with the majority of environmental impact stemming from device embodied carbon.

By \'Edouard Gu\'egain, Tristan Coignion
arXiv Machine Learning
Sep 14

A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS

The paper introduces a multi-vehicle dataset that includes camera, LiDAR, and radar sensor data along with scanned 3D models of all vehicles. Each vehicle’s pose and continuous kinematics are provided via RTK‑GNSS, enabling precise knowledge of the dynamic surroundings at any time. The dataset supports single‑ and multi‑object recordings with seven target vehicles, allowing evaluation of measurement effects such as occlusion and reflections thanks to known vehicle surface normals.

By Philipp Berthold, Bianca Forkel, Mirko Maehlisch
Simon Willison
Sep 12

Quoting Paul Ford

Simon Willison reflects on the evolving role of software developers in the age of AI, noting that while AI can produce high‑quality code, it also enables poor execution that leads to project failures. He argues that the industry is beginning to recognize the continued need for human collaboration and expertise to truly innovate. The piece highlights the tension between automation and the essential human element in software creation.

arXiv AI
Sep 12

A Mathematical Theory of Pragmatic Information

The paper introduces a pragmatic information theory that unifies communication, control, and decision-making through the isoteleia mapping, which formalizes equifinality by treating distinct semantic paths that lead to the same optimal action as pragmatically equivalent. It establishes a three-tier hierarchy of syntactic, semantic, and pragmatic information, defines pragmatic entropy, mutual information, channel capacity, and rate-distortion, and proves coding theorems that generalize Shannon’s results. The authors also present pragmatic value and cost of information, a Lagrangian dual framework for cross-layer optimization, and a pragmatic efficiency bound that quantifies the maximum net utility for resource-constrained intelligent systems, extending the theory to continuous messages and dynamic settings.

By Kai Niu, Ping Zhang
arXiv AI
Sep 12

ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies

arXiv:2609. 11697v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment.

By Jianming Ma, Rongjun Jin, Xiaxi Si, Yang Zhang, Yiheng Li, Yue Gao
arXiv AI
Sep 12

When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents

The paper introduces an admission‑audit protocol for continual embodied agents, arguing that update admission should balance error control with retained learning opportunities within a fixed interaction budget. It critiques a range‑based confidence gate for failing to certify unchanged old‑task behavior, and proposes a paired‑binomial construction that reduces this burden when outcome disagreements are rare. Experiments on a one‑step pushing diagnostic show that fresh paired checks admit a significant portion of updates while the range‑based gate admits none, and a learned‑dynamics stress test helps distinguish model bias from feedback‑selection error.

By Qinzhen Ma, Ruihai Wu
arXiv AI
Sep 12

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

The paper introduces 2AM, a system that separates memory and action execution in long‑horizon robot manipulation. 2AM stores task memory exclusively in a multimodal Agent, while a single RGB‑based, stateless Action Model performs motion based on language and optional 2D hints. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion—over 61 points higher than the best baseline—demonstrating that agent‑side memory and precise steering of the Action Model can substantially improve performance.

By Yutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry