arXiv Computer Vision By Waikit Xiu, Qiang Lu, Junbiao Chen, Xiying Li

PredErase: Training-Free Object-and-Effect Removal with Predictive Latent Guidance

Read the original on arXiv Computer Vision →

PredErase is a training‑free method that removes objects and their photometric effects by guiding a frozen Fill model with predictive latent cues. It expands the user‑provided mask to a contact‑band region, uses I‑JEPA to generate a context‑conditioned target for the hole, and aligns Fill’s completions with this target while keeping surrounding pixels fixed. On benchmarks such as RemovalBench, RORD‑Val, and DEFACTO‑Val, PredErase improves the native FLUX.2 backbone for instance‑only masks, though supervised removers still outperform it on full‑image metrics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 29

OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal

arXiv:2606. 28094v1 Announce Type: cross Abstract: Real-world object removal is challenging due to two key difficulties: the target object's non-local effects, such as shadows and reflections, which are difficult to model, and the fact that user-provided masks are often inaccurate or incomplete.

By Qinming Zhou, Chenxi Sun, Deyang Kong, Junhao He, Xiangheng Tang, Peike Yu, Haotian Wu, Leilei Cao, Linfeng Zhang
arXiv Computer Vision
Sep 7

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

DART is a new RGB‑D pretraining method for surgical vision foundation models that incorporates pseudo‑labeled depth maps as a pixel‑space reconstruction target during training. By adding a depth reconstruction head to DINOv2’s masked iBOT framework, DART improves representation quality without affecting downstream RGB‑only fine‑tuning or inference. Across eight surgical benchmarks—including segmentation, depth estimation, and image‑level recognition—DART outperforms both natural‑image and in‑domain baselines, demonstrating that geometric pseudo‑labels can strengthen foundation model pretraining without extra labels or inference cost.

By John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri
arXiv Computer Vision
4d ago

Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding

The paper introduces SGMA, a structure‑guided masked autoencoding framework designed for ultra‑high‑resolution scientific images. SGMA combines a content‑adaptive quadtree tokenizer that reduces gigapixel images to a fixed‑length sequence with a structure‑conditioned masking process that focuses reconstruction on spatially informative regions. The method, enhanced by Damped Accumulation to stabilize multi‑scale signals, achieves superior performance over standard MAE baselines on electron microscopy, whole‑slide optical microscopy, and X‑ray CT datasets, delivering significant accuracy gains and up to a 24.8× inference speedup.

By Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas, Amir Koushyar Ziabari, Xiao Wang, Peng Chen, Tao Luo, Toshio Endo, Fumiyoshi Shoji, Kento Sato, Kentaro Uesugi, Takayuki Nonoyama, Ryuji Kiyama, Masahiro Yoshida, Masaru Tezuka, Tetsuya Ishikawa, Satoshi Matsuoka, Masaharu Munetomo, Mohamed Wahib