arXiv Computer Vision

Depth-to-RGB: Repurposing a Frozen Depth Estimator for Geometry-Guided Compositing

Hugging Face Trending Papers
Jul 7

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets.

arXiv Computer Vision
Sep 7

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

DART is a new RGB‑D pretraining method for surgical vision foundation models that incorporates pseudo‑labeled depth maps as a pixel‑space reconstruction target during training. By adding a depth reconstruction head to DINOv2’s masked iBOT framework, DART improves representation quality without affecting downstream RGB‑only fine‑tuning or inference. Across eight surgical benchmarks—including segmentation, depth estimation, and image‑level recognition—DART outperforms both natural‑image and in‑domain baselines, demonstrating that geometric pseudo‑labels can strengthen foundation model pretraining without extra labels or inference cost.

By John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri
arXiv Computer Vision
1d ago

Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement

FreeInpaint is a feed‑forward 3D inpainting framework that reconstructs complete, geometrically consistent scenes directly from unposed multi‑view images with masked regions. It extends a 3D foundation model by propagating masked areas across views, using a Learnable Mask Attention mechanism to maintain reliable cross‑view correspondences and a Support Token Refinement strategy that injects diffusion‑generated auxiliary tokens for high‑fidelity completion. Experiments on diverse datasets show that FreeInpaint delivers superior inpainting quality without requiring pre‑computed camera poses, while maintaining fast inference speed.

By Jingyi Pan, Dan Xu, Qiong Luo
arXiv Computer Vision
Sep 28

Self-Supervised Perceptually Interpretable Monocular Depth Estimation

The paper introduces PIMDE, a self‑supervised monocular depth estimation framework that decomposes input images into perceptual feature maps, each encoding a specific visual cue. Separate depth branches process these maps to produce individual depth estimates, which are then fused explicitly. Experiments on the KITTI benchmark show that PIMDE matches the accuracy of existing self‑supervised methods while offering clearer insight into how each perceptual cue contributes to depth prediction.

By Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis