arXiv Computer Vision

PredErase: Training-Free Object-and-Effect Removal with Predictive Latent Guidance

PredErase is a training‑free method that removes objects and their photometric effects by guiding a frozen Fill model with predictive latent cues. It expands the user‑provided mask to a contact‑band region, uses I‑JEPA to generate a context‑conditioned target for the hole, and aligns Fill’s completions with this target while keeping surrounding pixels fixed. On benchmarks such as RemovalBench, RORD‑Val, and DEFACTO‑Val, PredErase improves the native FLUX.2 backbone for instance‑only masks, though supervised removers still outperform it on full‑image metrics.

arXiv AI
Jun 29

OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal

arXiv:2606. 28094v1 Announce Type: cross Abstract: Real-world object removal is challenging due to two key difficulties: the target object's non-local effects, such as shadows and reflections, which are difficult to model, and the fact that user-provided masks are often inaccurate or incomplete.

By Qinming Zhou, Chenxi Sun, Deyang Kong, Junhao He, Xiangheng Tang, Peike Yu, Haotian Wu, Leilei Cao, Linfeng Zhang
arXiv Computer Vision
Sep 7

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

DART is a new RGB‑D pretraining method for surgical vision foundation models that incorporates pseudo‑labeled depth maps as a pixel‑space reconstruction target during training. By adding a depth reconstruction head to DINOv2’s masked iBOT framework, DART improves representation quality without affecting downstream RGB‑only fine‑tuning or inference. Across eight surgical benchmarks—including segmentation, depth estimation, and image‑level recognition—DART outperforms both natural‑image and in‑domain baselines, demonstrating that geometric pseudo‑labels can strengthen foundation model pretraining without extra labels or inference cost.

By John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri
arXiv Computer Vision
4d ago

Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding

The paper introduces SGMA, a structure‑guided masked autoencoding framework designed for ultra‑high‑resolution scientific images. SGMA combines a content‑adaptive quadtree tokenizer that reduces gigapixel images to a fixed‑length sequence with a structure‑conditioned masking process that focuses reconstruction on spatially informative regions. The method, enhanced by Damped Accumulation to stabilize multi‑scale signals, achieves superior performance over standard MAE baselines on electron microscopy, whole‑slide optical microscopy, and X‑ray CT datasets, delivering significant accuracy gains and up to a 24.8× inference speedup.

By Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas, Amir Koushyar Ziabari, Xiao Wang, Peng Chen, Tao Luo, Toshio Endo, Fumiyoshi Shoji, Kento Sato, Kentaro Uesugi, Takayuki Nonoyama, Ryuji Kiyama, Masahiro Yoshida, Masaru Tezuka, Tetsuya Ishikawa, Satoshi Matsuoka, Masaharu Munetomo, Mohamed Wahib
arXiv Computer Vision
Sep 25

One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation

The paper introduces OVIE, a monocular novel-view synthesis method that eliminates the need for multi‑view training data. By using a frozen depth estimator to generate pseudo‑target views from single images and applying masked and adversarial losses, OVIE is trained on 30 million uncurated images. It achieves state‑of‑the‑art performance on RealEstate10K and DL3DV, produces highly consistent multi‑view trajectories, and runs at 116 FPS—over 600× faster than the fastest baseline.

By Adrien Ramanana Rahary, Nicolas Dufour, Patrick Perez, David Picard
arXiv Machine Learning
Aug 7

SR-JEPA: Learning Predictive Latent State in 3D Scenes

arXiv:2608. 05774v1 Announce Type: cross Abstract: Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce.

By Zihan Zhou, Qifu Wen, Xi Zeng
arXiv AI
Aug 28

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.

By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner