arXiv Computer Vision

$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction

arXiv AI
Jun 18

URDF Synthesis from RGB-D Sequences via Differentiable Joint Inference and Energy-Consistent Verification

arXiv:2606. 18861v1 Announce Type: cross Abstract: Reconstructing simulation-ready digital twins of articulated objects from sensor observations remains constrained by two persistent gaps: (i) part-level geometric reconstruction is decoupled from kinematic-parameter estimation, and (ii) the recovered models often violate basic dynamic invariants such as energy conservation, leading to drift when the URDF is replayed in physics simulators.

By Xinze Zhang
arXiv AI
Jun 29

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.

By Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
arXiv AI
Jun 2

The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

arXiv:2606. 01901v1 Announce Type: cross Abstract: We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image generator across multiple turns, making accumulated common ground directly observable as a rendered image.

By Sherzod Hakimov, Mattia D'Agostini, Ivan Samodelkin, David Schlangen
arXiv Computer Vision
2d ago

From Simulation to the Real-World: An In-Field 6D Pose Dataset and Baseline for Robotic Strawberry Harvesting

The paper presents the first real‑world 6D pose ground‑truth dataset for red‑stage strawberries, collected from 12,040 images at an actual farm using indirect camera pose recovery and 3D bounding‑box annotation. It also introduces a synthetic dataset rendered in NVIDIA Isaac Sim with scene‑level realism and domain randomization. Experiments show that models trained solely on synthetic data do not transfer well to in‑field images, but adding a small amount of real data significantly improves both translation and rotation accuracy across various backbone encoders.

By Woojung Son (Department of Agricultural and Biological Engineering, University of Florida), Won Suk Lee (Department of Agricultural and Biological Engineering, University of Florida), Zijing Huang (Department of Agricultural and Biological Engineering, University of Florida), Daeun Choi (Department of Agricultural and Biological Engineering, University of Florida), Catia Silva (Department of Electrical and Computer Engineering, University of Florida), Yu She (Edwardson School of Industrial Engineering, Purdue University), Yan Gu (School of Mechanical Engineering, Purdue University)
Hugging Face Trending Papers
Jun 22

Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI

Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predominantly designed for autonomous driving with vehicle-centric biases -- forward-facing cameras, far-field geometry, and static road priors -- limiting their applicability to embodied humanoid perception.

arXiv Machine Learning
Jun 19

Pose6DAug: Physically Plausible Multi-view Object Swapping for Robot Data Augmentation

arXiv:2606. 20118v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution.

By Jonghoon Lee, Seong Hyeon Park, Byungwoo Jeon, Minha Lee, Jinwoo Shin