ARGenSeg introduces an autoregressive generation-based approach for image segmentation that integrates seamlessly with multimodal large language models (MLLMs). Unlike prior methods that use boundary points or dedicated segmentation heads, ARGenSeg generates dense masks directly through visual token output and detokenization via a universal VQ‑VAE, enabling fine‑grained pixel‑level perception. The framework employs a next‑scale‑prediction strategy to parallelize token generation, resulting in faster inference while outperforming state‑of‑the‑art segmentation models on multiple datasets.
By Xiaolong Wang, Lixiang Ru, Ziyuan Huang, Kaixiang Ji, Dandan Zheng, Jingdong Chen, Jun Zhou
The paper introduces a two-stage framework for point-supervised change detection that leverages SAM2 priors to generate object-aware candidate masks and refines them with a lightweight CNN and uncertainty-aware loss. In the second stage, a teacher‑student self‑training loop with exponential moving average updates continuously improves pseudo‑labels and model performance. Experiments on WHU-CD, LEVIR-CD, and SYSU-CD show the method surpasses prior weakly supervised approaches and competes with fully supervised ones.
By Hailong Ning, Hao Wang, Yimeng Wang, Tao Lei, Renwei Dian, Asoke K. Nandi
The paper presents the first systematic evaluation of uncertainty quantification (UQ) methods applied to a foundation model for semantic segmentation. By fine‑tuning a lightweight DPT decoder on the pretrained SAM2 encoder, the authors benchmark four UQ approaches—Monte Carlo Dropout, Deep Sub‑Ensemble, Test‑Time Augmentation, and Evidential Deep Learning—across Cityscapes, NYUv2, and two out‑of‑domain settings, comparing segmentation accuracy, calibration, uncertainty quality, and inference time. The results reveal clear trade‑offs between predictive performance, reliability, and computational cost, underscoring both the promise and current limitations of uncertainty‑aware foundation models for real‑world deployment.
By Steven Landgraf, Joceline Hinz, Markus Ulrich
arXiv:2607. 05319v1 Announce Type: cross Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures.
By Rajat Rasal, Avinash Kori, Tian Xia, Ben Glocker
SAUF-Net is a semi‑supervised medical image segmentation framework that learns structure–appearance representations with uncertainty feedback. It decomposes bottleneck features into structural and appearance components, injects them into decoding, and uses auxiliary decoders and a dual‑head discriminator to estimate reliability and uncertainty. Experiments on ISIC‑2016 and Kvasir‑SEG show that SAUF‑Net surpasses state‑of‑the‑art methods, particularly when few labels are available.
By Qin Lu, Zheyang Jing, Yujie Yang, Jianwang Li, Chen Yi, Shaofeng Jiang
The paper introduces a two‑stage framework for point‑supervised change detection that leverages SAM2 priors to generate object‑aware candidate masks from sparse point annotations. In Stage I, a mask selection strategy converts generic segmentation outputs into reliable change pseudo‑labels, followed by a lightweight CNN refinement module with an uncertainty‑aware loss to enhance boundary quality. Stage II employs a teacher‑student self‑training loop, where the teacher is updated via exponential moving average and periodically refreshes pseudo‑labels, creating a closed‑loop optimization that alternates between pseudo‑label refinement and model re‑optimization. Experiments on WHU‑CD, LEVIR‑CD, and SYSU‑CD show the method surpasses prior weakly supervised approaches and competes with several fully supervised methods.