The paper presents a framework that builds a static point cloud prior map from past camera traversals, augmenting each point with DINOv3 semantic features. During runtime, a local prior patch is retrieved, encoded with a sparse voxel backbone, and fused with lifted multi‑view camera features in bird’s‑eye view. This fused representation is then used by sparse transformer heads to predict 3D objects and vectorized map elements, achieving improved performance on Argoverse 2 without requiring LiDAR for prior‑map construction or online inference.
By Markus K\"appeler, Rohit Mohan, Abhinav Valada
arXiv:2609.22353v1 Announce Type: new
Abstract: Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained p...
By Savio Cardoz, Santhiya Rajan
arXiv:2609.22813v1 Announce Type: cross
Abstract: We present \emph{commonsense ranked search} (CoRS), a novel path planner that turns an abstract instruction into a route that follows commonsense. Wh...
By Masafumi Endo, Kohei Honda, Ryo Yonetani
arXiv:2609.22974v1 Announce Type: cross
Abstract: In this paper, we investigate the Minimum Obstacle Displacement Planning problem from a robot motion planning perspective. The problem involves deter...
By Antony Thomas, Giulio Ferro, Fulvio Mastrogiovanni, Michela Robba, Marco Baglietto
arXiv:2609.23103v1 Announce Type: cross
Abstract: While simulation-ready deformable assets are essential for in-silico robotic manipulation tasks, existing generation frameworks typically assess phys...
By Guanxiong Chen, Yiduo Qu, Qianjun Xia, Pengyu Jing, Yixian Cheng, Bole Ma, Pengzhi Yang, Bingyang Zhou, Ziming Li, Shashwat Suri, Gongbo Sun, Chao Liu, Peter Yichen Chen, Ziqiu Zeng, Fan Shi
arXiv:2609.23910v1 Announce Type: cross
Abstract: Simulation-based evaluation provides a scalable and repeatable alternative to real-world evaluation of vision-language-action (VLA) policies. However...
By Xinyi Wang, Heng Hao, Wenjun Hu, Anna Enyu Li, Dizhi Ma, Karthik Ramani, Hankyu Moon, Yeong-Dae Kwon
arXiv:2609.23997v1 Announce Type: cross
Abstract: Multi-robot collaboration could enable more efficient and scalable solutions to complex robotic tasks, but collaboration under partial observability...
By Dorian Benhamou Goldfajn, Mason Nakamura, Saaduddin Mahmud, Justin Svegliato, Kyle H. Wray, Shlomo Zilberstein
arXiv:2609.24433v1 Announce Type: cross
Abstract: Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-t...
By Hung T. Ho, Khanh D. Nguyen, Quang D. Nguyen, Thanh Q. Duong, Ngan Le, Meng Guo, Vien A. Ngo, An T. Le
arXiv:2609.24660v2 Announce Type: cross
Abstract: Human demonstrations offer a scalable way to collect manipulation data, but their contacts may be unstable or infeasible when transferred to a robot...
By Shengcheng Luo, Xiaoyang Cheng, Hong Ying, Xiaoying Zhou, Jiaming Jiang, Haoran Guo, Wanlin Li, Ziyuan Jiao, Chenxi Xiao
arXiv:2609.24801v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation throu...
By Fernando Outeda, Gustavo Betarte, Juan Diego Campo, Fiorella Cravero
arXiv:2609.24906v1 Announce Type: cross
Abstract: Dormant tree pruning is labor-intensive yet essential for maintaining modern high-productivity fruit orchards. In this work, we focus on pruning of m...
By Abhinav Jain, Cindy Grimm, Stefan Lee
arXiv:2603.26687v2 Announce Type: replace-cross
Abstract: Hybrid aerial--ground robots can use thrust to cross obstacles that impede wheel-driven motion, but deciding how much thrust to apply during...
By Jiaxing Li, Ishaan Bhimwal, Wen Tian, Xinhang Xu, Junbin Yuan, Yuxin Guo, Sebastian Scherer, Muqing Cao
arXiv:2607.11498v2 Announce Type: replace-cross
Abstract: Vision-language-action (VLA) models require 3D spatial reasoning, yet RGB observations encode robot-object geometry only implicitly. Lifting...
By Byungkun Lee, Dongyoon Hwang, Dongjin Kim, Hojoon Lee, Hyunseung Kim, Jaegul Choo, Minho Park
arXiv:2609.25701v1 Announce Type: new
Abstract: We study distributed Byzantine-resilient actor-critic multi-agent reinforcement learning (AC-MARL), where agents collectively learn policies through lo...
By Haejoon Lee, Dimitra Panagou
arXiv:2609.25757v1 Announce Type: new
Abstract: What is the least recurrent memory needed to reproduce a specified expert under partial observability? The instantaneous requirement is the conditional...
By Xianyao Li, Fang Xu, Rui Min, Ruitong Tian, Jing Du
arXiv:2609.25351v1 Announce Type: cross
Abstract: We focus on human-robot collaborative transport, a challenging task of broad relevance spanning logistics, manufacturing, and the home, in which a us...
By Elvin Yang, Christoforos Mavrogiannis
arXiv:2609.25558v1 Announce Type: cross
Abstract: Vision-language-action policies benefit from geometric supervision, but current-frame geometry alone does not explicitly describe the changes associa...
By Jinu Pahk, Jesoon Kang, Taegeon Park, Jisu An, Soo Min Kimm, Jaejoon Kim, Byoung-Tak Zhang
arXiv:2406.04814v4 Announce Type: replace-cross
Abstract: Video diffusion models can enable embodied agents to anticipate plausible futures from the recent past, but they are typically trained offlin...
By Jason Yoo, Yingchen He, Saeid Naderiparizi, Dylan Green, Gido M. van de Ven, Geoff Pleiss, Frank Wood
arXiv:2609.26035v1 Announce Type: new
Abstract: Conversational agents often express answers in a uniformly confident register. We test whether expressed uncertainty, provenance-aware assertion, and e...
By Sebastian Cochinescu
arXiv:2609.24631v1 Announce Type: cross
Abstract: Autonomous parking in nonconvex and narrow environments remains challenging. Although optimal-control methods can explicitly enforce vehicle dynamics...
By Zhengbao Yao, Yuanfu Luo, Kehan Xue