ARGenSeg introduces an autoregressive generation-based approach for image segmentation that integrates seamlessly with multimodal large language models (MLLMs). Unlike prior methods that use boundary points or dedicated segmentation heads, ARGenSeg generates dense masks directly through visual token output and detokenization via a universal VQ‑VAE, enabling fine‑grained pixel‑level perception. The framework employs a next‑scale‑prediction strategy to parallelize token generation, resulting in faster inference while outperforming state‑of‑the‑art segmentation models on multiple datasets.
By Xiaolong Wang, Lixiang Ru, Ziyuan Huang, Kaixiang Ji, Dandan Zheng, Jingdong Chen, Jun Zhou
The paper introduces a two-stage framework for point-supervised change detection that leverages SAM2 priors to generate object-aware candidate masks and refines them with a lightweight CNN and uncertainty-aware loss. In the second stage, a teacher‑student self‑training loop with exponential moving average updates continuously improves pseudo‑labels and model performance. Experiments on WHU-CD, LEVIR-CD, and SYSU-CD show the method surpasses prior weakly supervised approaches and competes with fully supervised ones.
By Hailong Ning, Hao Wang, Yimeng Wang, Tao Lei, Renwei Dian, Asoke K. Nandi
The paper presents the first systematic evaluation of uncertainty quantification (UQ) methods applied to a foundation model for semantic segmentation. By fine‑tuning a lightweight DPT decoder on the pretrained SAM2 encoder, the authors benchmark four UQ approaches—Monte Carlo Dropout, Deep Sub‑Ensemble, Test‑Time Augmentation, and Evidential Deep Learning—across Cityscapes, NYUv2, and two out‑of‑domain settings, comparing segmentation accuracy, calibration, uncertainty quality, and inference time. The results reveal clear trade‑offs between predictive performance, reliability, and computational cost, underscoring both the promise and current limitations of uncertainty‑aware foundation models for real‑world deployment.
By Steven Landgraf, Joceline Hinz, Markus Ulrich
arXiv:2607. 05319v1 Announce Type: cross Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures.
By Rajat Rasal, Avinash Kori, Tian Xia, Ben Glocker
SAUF-Net is a semi‑supervised medical image segmentation framework that learns structure–appearance representations with uncertainty feedback. It decomposes bottleneck features into structural and appearance components, injects them into decoding, and uses auxiliary decoders and a dual‑head discriminator to estimate reliability and uncertainty. Experiments on ISIC‑2016 and Kvasir‑SEG show that SAUF‑Net surpasses state‑of‑the‑art methods, particularly when few labels are available.
By Qin Lu, Zheyang Jing, Yujie Yang, Jianwang Li, Chen Yi, Shaofeng Jiang
The paper introduces a two‑stage framework for point‑supervised change detection that leverages SAM2 priors to generate object‑aware candidate masks from sparse point annotations. In Stage I, a mask selection strategy converts generic segmentation outputs into reliable change pseudo‑labels, followed by a lightweight CNN refinement module with an uncertainty‑aware loss to enhance boundary quality. Stage II employs a teacher‑student self‑training loop, where the teacher is updated via exponential moving average and periodically refreshes pseudo‑labels, creating a closed‑loop optimization that alternates between pseudo‑label refinement and model re‑optimization. Experiments on WHU‑CD, LEVIR‑CD, and SYSU‑CD show the method surpasses prior weakly supervised approaches and competes with several fully supervised methods.
arXiv:2607. 05955v1 Announce Type: cross Abstract: Interactive 3D segmentation aims to extract object masks in point clouds with minimal user clicks.
By Shuheng Zhang, Feng Wu
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
By Nikolai R\"ohrich, Julian Glei{\ss}ner, Ahmed H. A. Ibrahim, Silvan Mertes, Tobias Huber
The paper presents the first systematic evaluation of uncertainty quantification (UQ) methods applied to a foundation model for semantic segmentation. By fine‑tuning a lightweight DPT decoder on the pretrained SAM2 encoder, the authors benchmark four UQ approaches—Monte Carlo Dropout, Deep Sub‑Ensemble, Test‑Time Augmentation, and Evidential Deep Learning—across Cityscapes, NYUv2, and two out‑of‑domain settings. The study compares segmentation accuracy, calibration, uncertainty quality, and inference time, revealing trade‑offs between predictive performance, reliability, and computational cost.
arXiv:2607.12896v3 Announce Type: replace
Abstract: Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fr...
By Yunzhou Li, Jiesi Hu, Yanwu Yang, Hanyang Peng, Chenfei Ye, Jianfeng Cao, Yixuan Yuan, Ting Ma
FoRIS is a training‑free in‑context segmentation framework that refines foreground masks through a coarse‑to‑fine process. It operates in three stages—Foreground Purification, Localization, and Consolidation—to suppress background noise, pinpoint target regions, and reconstruct complete foreground structures. The method achieves state‑of‑the‑art performance, improving mIoU by 4.5 and 4.8 points in 1‑shot and 5‑shot settings respectively.
By Ming Hu, Jianfu Yin, Mingyu Dou, Miaomiao Zhang, Yao Wang, Cong Hu, Bingliang Hu, Quan Wang
The paper introduces the Real‑Calibrated Synthetic‑First Data Engine, a modular pipeline that integrates controllable diffusion‑based synthetic image generation with multi‑stage curation, filtering, and optional uncertainty‑driven selection and human verification. Designed as a CLI‑based framework, it allows independent configuration of generation, filtering, selection, and validation modules to enhance reproducibility and flexibility in real‑world data workflows. Empirical tests on human pose estimation demonstrate that synthetic data can boost a real‑data baseline when used as low‑cost augmentation, though synthetic‑only training still lags behind real‑only performance, underscoring the importance of data‑centric orchestration in low‑data regimes.
By Yukang Shen, Zhiguo Liu, Yingshu Li, Yan Huang