arXiv:2609.23731v1 Announce Type: cross
Abstract: Robotic systems are typically composed of multiple independently developed modules that work together to perceive, predict, and act in the environmen...
By Rista Baral
arXiv:2509.08494v2 Announce Type: replace-cross
Abstract: As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures....
By Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes, Jacy Reese Anthis
arXiv:2609.24410v1 Announce Type: new
Abstract: Speech-to-text engines are extremely needed nowadays for different applications, representing an essential enabler in human-robot interaction. Still, s...
By Ali A. Safieh, Ibrahim Abu Alhaol, Rawan Ghnemat
arXiv:2609.23367v1 Announce Type: cross
Abstract: FORM is a domain-specific symbolic manipulation language widely used in particle physics for processing the very large algebraic expressions arising...
By Bakar Chargeishvili
arXiv:2609.22716v1 Announce Type: new
Abstract: Image-to-LiDAR registration estimates the camera pose of an image with respect to a LiDAR point cloud. It has diverse applications in autonomous drivin...
By Zijun Li, Xiaotian Sun, Xuelun Shen, Yao Dai, Sheng Ao, Yangyang Shi, Jakob Engel, Zhipeng Cai, Cheng Wang
arXiv:2609.22868v1 Announce Type: new
Abstract: End-to-end driving requires planning-relevant bird's-eye-view (BEV) representations, but existing pretraining approaches often rely on task annotations...
By Jaeha Song, Soonmin Hwang
arXiv:2609.23049v1 Announce Type: new
Abstract: Large-scale scene reconstruction is a critical foundational technology in robotic autonomous systems such as 3D mapping and autonomous driving. In rece...
By Haolin Yu, Jiadong Tang, YiXian Wang, Yu Gao, Shi He, Zhilin Lai, Yi Yang, Mengyin Fu
arXiv:2609.23561v1 Announce Type: new
Abstract: Deploying efficient neural networks is essential in resource-constrained environments, yet compact models often sacrifice interpretability - a critical...
By Aleks Czufarow, Ihor Babin
arXiv:2609.23673v1 Announce Type: new
Abstract: In vision-based off-road navigation, a robot needs to know not only which obstacles to avoid but also which terrain is better. The first is handled by...
By Ji-Hoon Hwang, Jisung Bae, E-In Son, Dong-Wook Kim, Jung-Taak Kim, Seung-Woo Seo
arXiv:2609.23961v1 Announce Type: new
Abstract: Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflectio...
By Zhihao Xing, Yingyu Wang, Liang Zhao, Shoudong Huang
arXiv:2609.24253v1 Announce Type: cross
Abstract: 3D Gaussian Splatting (3DGS) provides high-fidelity scenes for large-scale embodied simulation, but constructing large-scale urban assets remains con...
By Zhongrui You, Zhen Li, Junli Liu, Zhigang Wang, Bin Zhao
arXiv:2609.24976v1 Announce Type: cross
Abstract: Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple pre...
By Haoran Yuan, Zekai Wang, Boning Shao, Haoran Lu, Trevor Darrell, Ismini Lourentzou, Wei Zhan
arXiv:2603.11984v2 Announce Type: replace
Abstract: Diffusion-based visuomotor policies model complex action distributions through iterative denoising, but repeated inference adds latency to robotic...
By Chongyang Xu, Yixian Zou, Tianyu Yang, Fanman Meng, Ziliang Feng, Li Lu, Shuaicheng Liu
arXiv:2411.00527v5 Announce Type: replace-cross
Abstract: Utilizing the complementary strengths of wavelength-specific range or depth sensors is crucial for robust computer-assisted tasks such as aut...
By Vanessa Wirth, Johanna Br\"aunig, Nikolai Hofmann, Martin Vossiek, Tim Weyrich, Marc Stamminger
NAVIR is an end‑to‑end audio‑visual speech recognition system designed for the BrainChip Akida neuromorphic processor, which only supports sequential 2‑D convolutions. The architecture separates spatial and temporal encoding into three AkidaNet modules—per‑frame visual, temporal video, and spectrogram audio encoders—fused by a lightweight predictor and decoded with constrained beam search. Trained with CTC on noise‑augmented audio and fine‑tuned via quantization‑aware training, the quantized model achieves 14.0% WER on GRID’s unseen‑speaker split and 3.3% on overlapped‑speaker split, outperforming audio‑only baselines, and delivers 98.6% command accuracy at 1.5% WER on an industrial‑command corpus, while offering a 13‑fold energy advantage over conventional ANNs and roughly 5‑fold lower energy per inference than a Raspberry Pi CPU.
By Leonidas Delimpasis, Panagiota Moraiti, Antonis Porichis, Panos Chatzakos, Michail Karamousadakis
The paper presents a causal analysis of a compressed VLA policy that performs well in offline tests but fails in closed‑loop execution on a simulated pick‑and‑place task. An 8‑layer distillation of Octo‑Base retains most parameters and passes all offline metrics, yet collapses during deployment, with early stages degrading gradually and final transport failing entirely. The failure is traced to a negative, late‑heavy residual in the action trace, and standard remedies (continued training, offline data, command‑level compensation, clamping) do not restore performance; only a minimal‑pair intervention that mixes deployment‑distribution rollouts with teacher data restores parity with the teacher.
whyItMatters":"The study demonstrates that offline validation metrics alone are insufficient to guarantee closed‑loop success for compressed policies, highlighting the need for targeted deployment‑time testing and interventions."
By Fengze Jia (The Ohio State University)
Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization introduces U‑GROW, a lightweight sampling layer that directs more model rollouts toward states with high policy uncertainty, identified as decision‑sensitive stages where small action differences can alter task outcomes. By modifying only the branched‑start distribution, U‑GROW can be integrated into existing model‑based reinforcement learning pipelines without changing the policy optimization objective. Experiments on simulated and real‑world manipulation tasks demonstrate that U‑GROW improves the efficiency and effectiveness of policy optimization for Vision‑Language‑Action models.
By Yifei Sheng, Haoxiang Ren, Zhilong Zhang, Haonan Wang, Runjie Xu, Yihao Sun, Nan Tang, Zhichao Wu, Lei Yuan, Haoxin Lin, Yang Yu
arXiv:2609.22332v1 Announce Type: cross
Abstract: Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act...
By Jiadi You, Qize Yu, Yue Chen, Minghong Cai, Zhide Zhong, Yuran Wang, Bowen Ping, Jiaqi Liang, Zhenhao Shen, Haodong Yan, Yinchuan Li, Ruihai Wu, Xiaojuan Qi, Yingcong Chen
arXiv:2609.23492v1 Announce Type: new
Abstract: Perception for embodied agents is video-based, often multi-view (ego, exo, or both), and inherently continual, with simultaneous task and viewpoint shi...
By Hongwei Yan, Kanglei Zhou, Yuchen Liu, Qingyu Shi, Yi Zhong, Liyuan Wang
The paper introduces Action‑Slot, a structured action‑centric representation learning framework for multi‑agent atomic activity understanding. It reformulates slot attention into activity‑aligned slots, parallel spatio‑temporal updates, and background regularization to disentangle concurrent, asynchronous activities directly from raw video. Additionally, an attention‑difference pseudo‑mask method enables weakly supervised localization, and a new synthetic dataset, TACO, provides balanced atomic activity coverage with pixel‑level annotations.
By Yu-Ho Chang, Chi-Hsi Kung, Yi-Hsuan Tsai, Yi-Ting Chen