arXiv:2609.13851v1 Announce Type: cross
Abstract: Post-training vision-language-action (VLA) models for specific robots and tasks requires in-domain demonstrations, yet collecting diverse robot data...
By Chenwei Wang, Dianye Huang, Match W. L. Ko, Chenjia Bai, Zhongliang Jiang
arXiv:2609.13733v1 Announce Type: new
Abstract: Stable and reliable 4D spatial understanding is fundamental for autonomous driving systems. While feedforward reconstruction networks can estimate came...
By Meng-Li Shih, Shih-Yang Su, Yuliang Zou, Hao Xiang, Haidong Zhu, Vincent Casser, Brian Curless, Dmitry Kalenichenko, Mingxing Tan, Dragomir Anguelov
Stereo4DWalker is a 4D-aware embodied navigation model that uses stereo video inputs to explicitly construct structured representations of geometry and motion. These 4D structures are incorporated into a navigation transformer via 4D-conditioned attention layers, enabling the agent to learn robust urban navigation. The authors also curate a large-scale stereo navigation dataset with automatically annotated actions from Internet stereo videos, and demonstrate that Stereo4DWalker outperforms state‑of‑the‑art methods while requiring only 1.5% of the training data.
By Wentao Zhou, Xuweiyi Chen, Vignesh Rajagopal, Jeffrey Chen, Rohan Chandra, Zezhou Cheng
arXiv:2609.14615v1 Announce Type: cross
Abstract: Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world env...
By Guocun Wang, Kenkun Liu, Guorui Song, Jing Lin, Zhe Huang, Luyuan Zhang, Dake Zhong, Choo Sin Wai, Xiaoguang Han, Haoqian Wang
arXiv:2609.14567v1 Announce Type: cross
Abstract: Reinforcement learning (RL) has shown considerable promise for robotic decision-making, yet deploying multi-agent RL (MARL) on physical multi-robot s...
By Abdalwhab Bakheet Mohamed Abdalwhab, Giovanni Beltrame, David St-Onge
arXiv:2609.14899v1 Announce Type: new
Abstract: Neural 3D scene editing is often evaluated by semantic alignment alone, although a convincing result may alter unrelated content or become inconsistent...
By Sariah Patro, Arjun Mehra, Nikhil Bhatia
arXiv:2609.14313v1 Announce Type: cross
Abstract: Robust point tracking in endoscopic videos is essential for computer-assisted intervention and autonomous robotic surgery, enabling continuous regist...
By Jiaming Zhang, Zijian Wu, Mehran Armand, Septimiu Salcudean
arXiv:2609.13295v1 Announce Type: cross
Abstract: Diffusion learning leverages the statistical mechanism of diffusion processes for learning, reasoning, and inferring complex distributions from data....
By Max Muchen Sun, Cem Bilaloglu, Ananya Rao, Stefan Ivic, Guillaume Sartoretti, Kathleen Fitzsimons, Ian Abraham, Sylvain Calinon, Todd Murphey
arXiv:2609.14261v1 Announce Type: cross
Abstract: Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative...
By Prajwal Koirala, Mark Campbell
arXiv:2609.13352v1 Announce Type: cross
Abstract: We present three robotic art installations which explore the aesthetics of adaptive behavior. Through embodied machine leaning and digital evolution,...
By Sofian Audry, Stephen Kelly
arXiv:2609.13308v1 Announce Type: cross
Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...
By Sarthak Sattigeri
arXiv:2609.13243v1 Announce Type: cross
Abstract: We present GzDRL, a novel single-process reinforcement learning (RL) framework for Gazebo that overcomes longstanding bottlenecks in scalable, reprod...
By Amal Dev Haridevan, Junjie Kang, Jinjun Shan
arXiv:2609.13606v1 Announce Type: cross
Abstract: Multi-arm robotic harvesting offers a promising path to improve harvesting efficiency and reduce reliance on manual labor. However, practical deploym...
By Vrishan Inukollu, Adyan Zaman, Anvi Kudaraya, Carlos Lazcano, Yuankai Zhu, Stavros Vougioukas, Xiaofan Yu
arXiv:2609.14802v1 Announce Type: cross
Abstract: Importance weights are essential in domain adaptation under label shift, yet their utility is often undermined by the finite sample uncertainty assoc...
By Mushan Li, Kihyun Han, Yanyuan Ma
arXiv:2609.13225v1 Announce Type: cross
Abstract: Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fail...
By Sarthak Sattigeri
arXiv:2606.20980v2 Announce Type: replace-cross
Abstract: As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action mod...
By Adrian Cespedes, Marcelo Chincha, Dunant Cusipuma, Victor Flores-Benites, David Ortega, Arturo Deza
arXiv:2609.15726v1 Announce Type: cross
Abstract: Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not conve...
By Zhenjie Yang, Yideng Zhang, Dongjie Zhang, Chenyu Jiang, Xianshuai Liu, Yufeng Li, Zuhao Ge, Xingyu Jiao, Zheng Zhang, Kaiyu He, He Wang, Yuwen Zhong, Yi Deng, Muyun Jiang, Xianliang Huang, Haisheng Su, Donghang Zhang, Jian Zhang, Xue Yang, Hongyang Li, Zuxuan Wu, Yu-Gang Jiang, Xiaosong Jia, Junchi Yan
arXiv:2609.15169v1 Announce Type: new
Abstract: Driving vision-language-action (VLA) models increasingly reason before acting, but their intermediate reasoning is often weakly grounded in physical sc...
By Xiao Liu, Haoyu Li, Jianghao Leng, Lin Wang, Chao Sun
arXiv:2609.13232v1 Announce Type: new
Abstract: A robot pruning trees needs two facts per pixel: whether it belongs to a tree, and its distance. Both are usually obtained via task heads attached to a...
By Yida Lin, Bing Xue, Mengjie Zhang, Sam Schofield, Richard Green
The paper introduces a fast Bayesian method for estimating homographies from noisy point correspondences, providing a posterior distribution over the homography parameters. A closed‑form solution for the posterior mean in homogeneous coordinates is derived, complemented by an iterative Bayesian approach to address non‑linearities. Experiments on synthetic data and real image stitching show improved accuracy over DLT and supply uncertainty estimates for the homography.
By Hanne Beuter, Sebastian Dorn