Hugging Face Blog

Pollen-Vision: Unified interface for Zero-Shot vision models in robotics

arXiv AI
Jul 1

MVP-Nav: Multi-layer Value Map Planner Navigator

arXiv:2606. 31919v1 Announce Type: cross Abstract: Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment.

By Wenyuan Xie, Shaokai Wu, Yijin Zhou, Yanbiao Ji, Guodong Zhang, Bayram Bayramli, Qiuchang Li, Xunchu Zhou, Yue Ding, Hongtao Lu
arXiv Computer Vision
2d ago

From Simulation to the Real-World: An In-Field 6D Pose Dataset and Baseline for Robotic Strawberry Harvesting

The paper presents the first real‑world 6D pose ground‑truth dataset for red‑stage strawberries, collected from 12,040 images at an actual farm using indirect camera pose recovery and 3D bounding‑box annotation. It also introduces a synthetic dataset rendered in NVIDIA Isaac Sim with scene‑level realism and domain randomization. Experiments show that models trained solely on synthetic data do not transfer well to in‑field images, but adding a small amount of real data significantly improves both translation and rotation accuracy across various backbone encoders.

By Woojung Son (Department of Agricultural and Biological Engineering, University of Florida), Won Suk Lee (Department of Agricultural and Biological Engineering, University of Florida), Zijing Huang (Department of Agricultural and Biological Engineering, University of Florida), Daeun Choi (Department of Agricultural and Biological Engineering, University of Florida), Catia Silva (Department of Electrical and Computer Engineering, University of Florida), Yu She (Edwardson School of Industrial Engineering, Purdue University), Yan Gu (School of Mechanical Engineering, Purdue University)
arXiv AI
Jun 19

Human Universal Grasping

arXiv:2606. 17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality.

By Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
Hugging Face Trending Papers
Jul 6

From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model

Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing view-robust Vision-Language-Action (VLA) policies tolerate such camera variations only when the camera extrinsics are explicitly provided, making them fragile and hard to use especially when view robustness is critical.

arXiv AI
Jun 15

Schr\"odinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation

arXiv:2512. 21201v3 Announce Type: replace-cross Abstract: Zero-shot object navigation (ZSON) requires robots to find target objects in unseen environments without task-specific fine-tuning or pre-built maps, a key capability for general-purpose service robots.

By Yu He, Da Huang, Zhenyang Liu, Zixiao Gu, Qiang Sun, Guangnan Ye, Yanwei Fu, Yu-Gang Jiang
arXiv Machine Learning
Aug 7

Dynamic Object Masks as Goal Representations for Visual Goal-Conditioned Reinforcement Learning

arXiv:2510. 06277v2 Announce Type: replace-cross Abstract: Goal-conditioned reinforcement learning (GCRL) offers a unified way to pursue diverse tasks, yet most existing methods rely on state- or position-based goal representations that are unavailable in real-world robotics.

By Fahim Shahriar, Cheryl Wang, Alireza Azimi, Gautham Vasan, Hany Hamed, Abhishek Naik, A. Rupam Mahmood, Colin Bellinger
Hugging Face Trending Papers
Jul 24

Visual Relocalization from Sparse Views in Aliased and Low-Texture Environments via Novel View Synthesis

Visual localization becomes extremely challenging in planetary-like terrains characterized by low texture, perceptual aliasing, harsh illumination, and sparse, weakly overlapping viewpoints induced by forward rover motion and unconstrained driving directions. Under these conditions, state-of-the-art image-to-image and image-to-map matching pipelines suffer significant performance degradation.