Autonomous endoscopic navigation requires the policy model to predict actions from texture-poor monocular observations, make safe control decisions, and retain evidence of lesions after they leave the...
arXiv:2606. 29247v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics.
By Jiashuo Sun, Yue He, Wenxuan Liu, Tao Mao, Jiazheng Wang, Xiang Chen, Min Liu
arXiv:2606. 17340v1 Announce Type: cross Abstract: Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment.
By Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath
The paper introduces 3DWay, a method that predicts 3D-consistent waypoints for robot manipulation by first generating multi‑view consistent 2D waypoints and then triangulating them. This approach addresses the 3D ambiguity inherent in 2D trajectory predictions and leverages pretrained vision‑language models to provide explicit 3D motion specifications. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing generalization for robot manipulation tasks.
By Ziqin Huang, Yingyue Li, Chenyangguang Zhang, Ruida Zhang, Yuxin Chen, Gu Wang, Xingyu Liu, Masayoshi Tomizuka, Xiangyang Ji
The paper introduces 3DWay, a method that predicts 3D consistent waypoints for robot manipulation using multi‑view images. By first generating 2D waypoints that are consistent across views and then triangulating them, the approach provides explicit 3D motion specifications while leveraging pretrained vision‑language models. Experiments demonstrate that 3DWay improves 3D spatial grounding and vision‑language reasoning, enhancing the generalization of robot manipulation policies.
arXiv:2608. 11204v1 Announce Type: cross Abstract: Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.
By Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang
FOCAL‑VLA is a framework that improves vision‑language‑action models by combining subtask‑guided geometry distillation with implicit world modeling. It transfers geometric knowledge from VGGT to focus on subtask‑relevant image regions and uses Track4World features to capture future 3D evolution, guiding action generation without running these models at inference time. Experiments demonstrate that FOCAL‑VLA outperforms baselines on both simulation benchmarks and real‑world manipulation tasks.
By Zhiyuan Gao, Di Wen, Yanxiang Zhan, Mohammad Khoshnazar, Jeroen Sch\"afer, Kunyu Peng, Michael Beetz
GaussVLA is a Vision‑Language‑Action model that enhances spatial reasoning by converting flat 2D visual tokens into compact 3D Gaussian tokens using a Gaussian Spatial Tokenizer. It further employs a Depth‑Aware Chain‑of‑Thought module to perform structured, non‑autoregressive geometric reasoning conditioned on language and flow‑time. In both simulated and real‑world tests, GaussVLA achieves high spatial‑manipulation success rates—93.5% on LIBERO and 100% on the Spatial suite—while using only 200 M parameters, outperforming SpatialVLA by 19.7% relative success.
By Md Selim Sarowar, Md Tanvir Islam, Sungho Kim, Sangtae Ahn
arXiv:2509.18778v2 Announce Type: replace-cross
Abstract: Visual imitation learning frameworks allow robots to learn manipulation skills from expert demonstrations. While existing approaches mainly f...
By Shijia Ge, Yijun Liu, Yinxin Zhang, Shuzhao Xie, Weixiang Zhang, Mingcai Zhou, Zhi Wang
arXiv:2606. 00095v1 Announce Type: cross Abstract: Vision-Language Navigation (VLN) enables embodied agents to reach target locations in unseen environments by following language instructions.
By Kailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu, Jingyu Gong, Xiaoling Wang, Liang He
SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.
By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
arXiv:2609.17021v1 Announce Type: new
Abstract: Autonomous wheel-loader control requires joint reasoning over task semantics, egocentric vision, proprioception, and 3D scene geometry. We present sens...
By Gopi Krishna Erabati, Bjarne Johannsen, Angus Stewart, Vardeep Singh Sandhu