arXiv:2608.31052v1 Announce Type: cross
Abstract: Semantic segmentation decomposes an image into distinct mask regions corresponding to different object categories, such as people, cars, signs or bui...
By Keith G. Mills, Evan B. Sanders, Gregory J. Matthews, Juliet K. Brophy
arXiv:2602. 10045v2 Announce Type: replace-cross Abstract: Current instance segmentation models achieve high performance on average predictions, but lack principled uncertainty quantification: their outputs are not calibrated, and there is no guarantee that a predicted mask is close to the ground truth.
By Kerri Lu, Dan M. Kluger, Stephen Bates, Sherrie Wang
The paper investigates how the CutMix data augmentation technique affects reliability and robustness in semantic segmentation. It evaluates two architectures—CNN-based DeepLabV3+ and transformer-based SegFormer—on both in-domain and out-of-domain data. Results show that CutMix has a minor effect on segmentation accuracy but consistently improves reliability, especially under distribution shifts, by enhancing calibration and uncertainty quality.
By Steven Landgraf, Markus Ulrich
arXiv:2605. 13674v2 Announce Type: replace-cross Abstract: Weakly supervised semantic segmentation (WSSS) trains dense pixel-level segmentation models from partial or coarse annotations such as bounding boxes, scribbles, or image-level tags.
By Stefano Colamonaco, Andrei-Bogdan Florea, Jaron Maene
The paper introduces a framework that adapts a diffusion model to a target urban domain using only imperfect pseudo‑labels, enabling the generation of high‑fidelity, target‑aligned images from semantic maps of any synthetic dataset. By filtering poor generations, correcting image‑label misalignments, and standardising semantics, the method transforms low‑effort synthetic data into competitive real‑domain training sets. Experiments on five synthetic and two real datasets show up to +8.0 %pt mIoU improvement over state‑of‑the‑art translation methods, demonstrating that rapidly constructed synthetic datasets can match the performance of high‑effort, manually designed ones.
By Damjan Kal\v{s}an, Denis Zavadski, Tim K\"uchler, Haebom Lee, Stefan Roth, Carsten Rother
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
By Nikolai R\"ohrich, Julian Glei{\ss}ner, Ahmed H. A. Ibrahim, Silvan Mertes, Tobias Huber
arXiv:2606. 14754v1 Announce Type: cross Abstract: Images can be segmented based on visual cues (i.
By Aviad Cohen Zada, Nadav Orenstein, Shai Avidan, Gal Oren
FoRIS is a training‑free in‑context segmentation framework that refines foreground masks through a coarse‑to‑fine process. It operates in three stages—Foreground Purification, Localization, and Consolidation—to suppress background noise, pinpoint target regions, and reconstruct complete foreground structures. The method achieves state‑of‑the‑art performance, improving mIoU by 4.5 and 4.8 points in 1‑shot and 5‑shot settings respectively.
By Ming Hu, Jianfu Yin, Mingyu Dou, Miaomiao Zhang, Yao Wang, Cong Hu, Bingliang Hu, Quan Wang
The paper presents the first systematic evaluation of uncertainty quantification (UQ) methods applied to a foundation model for semantic segmentation. By fine‑tuning a lightweight DPT decoder on the pretrained SAM2 encoder, the authors benchmark four UQ approaches—Monte Carlo Dropout, Deep Sub‑Ensemble, Test‑Time Augmentation, and Evidential Deep Learning—across Cityscapes, NYUv2, and two out‑of‑domain settings, comparing segmentation accuracy, calibration, uncertainty quality, and inference time. The results reveal clear trade‑offs between predictive performance, reliability, and computational cost, underscoring both the promise and current limitations of uncertainty‑aware foundation models for real‑world deployment.
By Steven Landgraf, Joceline Hinz, Markus Ulrich
SegCol is a new dataset and benchmark for semantic segmentation of colon fold edges and surgical instruments in colonoscopy images, derived from the EndoMapper dataset. It offers manually annotated pixel‑level masks for three instrument classes and thin fold‑edge structures across temporally consistent image sequences, and serves as the basis for the SegCol Challenge within the EndoVis Challenge at MICCAI 2024. The study evaluates supervised segmentation and annotation‑efficient active learning, analyzes various segmentation metrics under structural perturbations, and highlights how metric behavior depends on target structure, underscoring the need for carefully selected evaluation protocols in endoscopic segmentation.
By Xinwei Ju, Rema Daher, Razvan Caramalau, Baoru Huang, Danail Stoyanov, Francisco Vasconcelos
MedSAM-3 is a text‑promptable medical segmentation model that builds on the Segment Anything Model (SAM) by fine‑tuning it with medical images and semantic concept labels. It enables precise anatomical segmentation through open‑vocabulary text descriptions, moving beyond purely geometric prompts. The accompanying MedSAM-3 Agent incorporates multimodal large language models to perform complex reasoning and iterative refinement, and experiments across X‑ray, MRI, ultrasound, CT, and video modalities show it outperforms existing specialist and foundation models.
By Anglin Liu, Xu R. Cao, Yifan Shen, Yi Lu, Xiang Li, Qianqian Chen, Jintai Chen