arXiv Computer Vision By Andr\'e Amorim, Pedro F. Proen\c{c}a

AnalogDepth: Multi-view Geometry from FPV drones under Analog Video Transmission

Read the original on arXiv Computer Vision →

AnalogDepth adapts the Depth Anything 3 (DA3) visual geometry model to the spatially structured noise of analog video transmission (VTX) used in FPV drones. The method employs a parameter‑efficient training pipeline that uses student‑teacher knowledge distillation with Low‑Rank Adaptation (LoRA) on the DINOv2 backbone, and trains on a noise bank built from real FPV recordings rather than synthetic Gaussian noise. Experiments on six real FPV flight sequences across three indoor scenes show that this real‑noise training consistently lowers per‑frame depth RMSE and 3D reconstruction Chamfer distance compared to the pretrained DA3 baseline and Gaussian noise baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 23

DIFTA-3D: Depth-Consistent Instance-Level Feature Transfer and Adaptation of DINOv3 for 3D Detection

The paper introduces DIFTA-3D, a method that replaces the task‑specific visual branch in IIFNet3D with a frozen DINOv3 foundation model for RGB‑D 3D instance detection. It employs a depth‑consistent feature pipeline that projects points into calibrated RGB‑D frames, filters features with a metric depth‑residual check, caches accepted DINOv3 features, and aggregates them within proposal‑aligned RoI grids. Extensive experiments on ScanNetV2 show that the DINOv3 control achieves mAP scores of 76.15/60.93 at IoU thresholds 0.25/0.50, while the Conservative VAID recipe improves these to 76.59/62.16, indicating a modest gain from the proposed transfer recipe.

By Linman Wang, ZiFei Zhang, Chunran Zheng, Xiwang Dong, Jiarong Lin
arXiv Computer Vision
Sep 25

One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation

The paper introduces OVIE, a monocular novel-view synthesis method that eliminates the need for multi‑view training data. By using a frozen depth estimator to generate pseudo‑target views from single images and applying masked and adversarial losses, OVIE is trained on 30 million uncurated images. It achieves state‑of‑the‑art performance on RealEstate10K and DL3DV, produces highly consistent multi‑view trajectories, and runs at 116 FPS—over 600× faster than the fastest baseline.

By Adrien Ramanana Rahary, Nicolas Dufour, Patrick Perez, David Picard
arXiv Computer Vision
Sep 3

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

The paper introduces Self-Geometry, a plug‑and‑play test‑time adaptation framework that enforces explicit multi‑view geometric constraints on Vision Foundation Models (VFMs) using 2D pixel correspondences as pseudo ground truth. It combines Geometric Disentanglement Optimization—mixing Multi‑View and Epipolar Consistency losses with Gradient Disentanglement—to avoid gradient conflicts, a Frame Angular‑Neighbor sampler based on SO(3) geodesic distances to select informative views, and a Lightweight TTA module that adapts VFMs via LoRA. Experiments on six VFMs and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom) show consistent improvements in pose and geometry estimation.

By Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh