arXiv:2608.23354v1 Announce Type: cross
Abstract: Autonomous indoor navigation requires both semantic understanding and precise geometric control. We propose OptiSight, a hybrid framework that combin...
By Alperen Avan, Jordi Sanchez-Riera
arXiv:2607. 17559v1 Announce Type: cross Abstract: The Contrastive Olfaction-Language-Image Pre-training 2 (COLIP-2) model is a multimodal embeddings space that places olfaction as a first-class citizen among vision and language.
By Kordel Kade France
arXiv:2606. 31919v1 Announce Type: cross Abstract: Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment.
By Wenyuan Xie, Shaokai Wu, Yijin Zhou, Yanbiao Ji, Guodong Zhang, Bayram Bayramli, Qiuchang Li, Xunchu Zhou, Yue Ding, Hongtao Lu
The paper presents the first real‑world 6D pose ground‑truth dataset for red‑stage strawberries, collected from 12,040 images at an actual farm using indirect camera pose recovery and 3D bounding‑box annotation. It also introduces a synthetic dataset rendered in NVIDIA Isaac Sim with scene‑level realism and domain randomization. Experiments show that models trained solely on synthetic data do not transfer well to in‑field images, but adding a small amount of real data significantly improves both translation and rotation accuracy across various backbone encoders.
By Woojung Son (Department of Agricultural and Biological Engineering, University of Florida), Won Suk Lee (Department of Agricultural and Biological Engineering, University of Florida), Zijing Huang (Department of Agricultural and Biological Engineering, University of Florida), Daeun Choi (Department of Agricultural and Biological Engineering, University of Florida), Catia Silva (Department of Electrical and Computer Engineering, University of Florida), Yu She (Edwardson School of Industrial Engineering, Purdue University), Yan Gu (School of Mechanical Engineering, Purdue University)
arXiv:2606. 17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality.
By Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
arXiv:2603. 25937v2 Announce Type: replace-cross Abstract: Visual Navigation Models (VNMs) promise generalizable, robot navigation by learning from large-scale visual demonstrations.
By Maeva Guerrier, Karthik Soma, Jana Pavlasek, Giovanni Beltrame
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing view-robust Vision-Language-Action (VLA) policies tolerate such camera variations only when the camera extrinsics are explicitly provided, making them fragile and hard to use especially when view robustness is critical.
arXiv:2512. 21201v3 Announce Type: replace-cross Abstract: Zero-shot object navigation (ZSON) requires robots to find target objects in unseen environments without task-specific fine-tuning or pre-built maps, a key capability for general-purpose service robots.
By Yu He, Da Huang, Zhenyang Liu, Zixiao Gu, Qiang Sun, Guangnan Ye, Yanwei Fu, Yu-Gang Jiang
arXiv:2606. 30696v1 Announce Type: cross Abstract: Enabling robots to follow natural language commands to complete zero-shot long-horizon tasks remains challenging.
By Kaier Liang, Hengde Dai, Cristian-Ioan Vasile
arXiv:2506. 11585v2 Announce Type: replace-cross Abstract: We introduce OV-MAP, a novel approach to open-world 3D mapping for mobile robots by integrating open-features into 3D maps to enhance object recognition capabilities.
By Juno Kim, Yesol Park, Hye-Jung Yoon, Byoung-Tak Zhang
arXiv:2510. 06277v2 Announce Type: replace-cross Abstract: Goal-conditioned reinforcement learning (GCRL) offers a unified way to pursue diverse tasks, yet most existing methods rely on state- or position-based goal representations that are unavailable in real-world robotics.
By Fahim Shahriar, Cheryl Wang, Alireza Azimi, Gautham Vasan, Hany Hamed, Abhishek Naik, A. Rupam Mahmood, Colin Bellinger
Visual localization becomes extremely challenging in planetary-like terrains characterized by low texture, perceptual aliasing, harsh illumination, and sparse, weakly overlapping viewpoints induced by forward rover motion and unconstrained driving directions. Under these conditions, state-of-the-art image-to-image and image-to-map matching pipelines suffer significant performance degradation.