InfiNoVA is a data‑augmentation framework that transforms synchronized multi‑camera demonstrations into a dense, geometrically consistent set of training views by reconstructing each manipulation trajectory as a time‑varying 3D Gaussian. The method renders novel observations from sampled camera poses while preserving the original state‑action pairs, improving frame‑level fidelity and temporal consistency compared to generative synthesis. Across four real‑world manipulation tasks, policies trained with InfiNoVA achieve 5.4× higher average success under unseen randomized viewpoints than VISTA‑based augmentation and 1.7× higher success than training on all five physical camera views.
By Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave
arXiv:2606. 10862v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task-relevant objects are fully visible.
By Taishan Li, Jiwen Zhang, Siyuan Wang, Xuanjing Huang, Zhongyu Wei
arXiv:2606. 31585v1 Announce Type: cross Abstract: The remarkable scalability of Transformers has expanded their application to 3D computer vision, where camera-aware positional encoding is crucial for providing spatial cues in multi-view geometry.
By Shun Kenney, Teppei Suzuki
arXiv:2606. 20118v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution.
By Jonghoon Lee, Seong Hyeon Park, Byungwoo Jeon, Minha Lee, Jinwoo Shin
arXiv:2511.16030v3 Announce Type: replace
Abstract: 3D Gaussian Splatting (3DGS) enables efficient, high-fidelity novel view synthesis, yet its performance degrades severely under sparse-view supervi...
By Zijian Wu, Mingfeng Jiang, Zidian Lin, Ying Song, Ziqian Lu, Qun Wu, Hanjie Ma
arXiv:2603.09632v5 Announce Type: replace-cross
Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, subsequently extending into numerous spatial AI ap...
By Yueen Ma, Zenglin Xu, Irwin King
arXiv:2610.07958v1 Announce Type: new
Abstract: Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across...
By Minhyeok Lee, Jungho Lee, Minseok Kang, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee
arXiv:2607. 11498v1 Announce Type: cross Abstract: Vision-language-action (VLA) models predict robot actions from visual observations and language instructions.
By Byungkun Lee, Dongyoon Hwang, Dongjin Kim, Hojoon Lee, Minho Park, Jaegul Choo
arXiv:2606. 24472v1 Announce Type: cross Abstract: Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their visual tokens remain grounded in 2D image coordinates rather than the calibrated geometry of the robot's cameras -- a mismatch especially pronounced in multi-camera setups, where views are coupled by known intrinsics and extrinsics yet processed as independent images.
By Yue Peng, Yongzhe Zhao, Artur Habuda, Khuyen Pham, Yanheng Zhu, Tran Nguyen Le, Fares Abu-Dakka, Li Guo
arXiv:2608.21402v1 Announce Type: cross
Abstract: World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera...
By Bingqi Huang, Bingchuan Wei, Yingkai Cai, Zhaokui Wang
Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set of context views and process them jointly. This limits their applicability to online scenarios where calibrated views arrive sequentially and the scene must be updated causally.
arXiv:2607. 00832v1 Announce Type: cross Abstract: A single panorama captures the full visual sphere from one camera center, yet confines users to looking around in place without enabling true scene exploration.
By Zhenjia Li, Jinrang Jia, Yifeng Shi