Universal Image Segmentation with Mask2Former and OneFormer
Related stories
Introducing Segment Anything: Working toward the first foundation model for image segmentation
LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter
arXiv:2607. 00687v1 Announce Type: cross Abstract: Comparing transformer backbones for image segmentation is confounded: each is paired with a different decoder, recipe, and pretraining, so reported differences rarely reflect the backbone itself.
Segmentation of Bovid Dentition Under Imperfect Annotations: A Comparative Study of Convolutional and Attention Models
arXiv:2608.31052v1 Announce Type: cross Abstract: Semantic segmentation decomposes an image into distinct mask regions corresponding to different object categories, such as people, cars, signs or bui...
Mask Proposal Voting Based on Geodesic Framework for Robust Image Segmentation
arXiv:2606. 14912v1 Announce Type: cross Abstract: Despite great advances, finding accurate segmentation remains a challenging task, especially in scenarios with cluttered backgrounds, complex intensity variations and topology appearance.
A Comprehensive Survey of Medical Image Segmentation: Challenges, Benchmarks, and Beyond
arXiv:2606. 16153v1 Announce Type: cross Abstract: Medical image segmentation plays a critical role in clinical diagnostics, treatment planning, disease monitoring, and neurological disorder identification.
Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
arXiv:2605.18177v2 Announce Type: replace Abstract: Query-based Vision Transformer segmentation models typically reconstruct dense spatial feature maps to predict masks, inheriting design patterns fr...
Conformal Prediction Sets for Instance Segmentation
arXiv:2602. 10045v2 Announce Type: replace-cross Abstract: Current instance segmentation models achieve high performance on average predictions, but lack principled uncertainty quantification: their outputs are not calibrated, and there is no guarantee that a predicted mask is close to the ground truth.
Lightweight Vision Transformer-Based U-Net for Brain Tumor Segmentation from MRI
The paper introduces a lightweight Vision Transformer‑based U‑Net for brain tumor segmentation from MRI, combining U‑Net’s hierarchical feature extraction with a compact ViT bottleneck to capture both local and global context. With only 2.6 million trainable parameters, the model achieves a mean Intersection over Union of 0.8100 and a Dice score of 0.8446 on the TCGA LGG dataset, surpassing the baseline U‑Net by 3.75% and 3.15% respectively. Extensive quantitative and qualitative analyses, including confusion matrices, precision‑recall curves, and tumor size dependency studies, demonstrate the method’s effectiveness and robustness.
FSANet: Frequency-Spatial Aware Network for Image Segmentation
Image segmentation remains challenging due to occlusions, poor lighting, and irregular structures. Although transformer-based methods achieve high accuracy, they rely heavily on long-range spatial fea...
FSANet: Frequency-Spatial Aware Network for Image Segmentation
arXiv:2609.16773v1 Announce Type: new Abstract: Image segmentation remains challenging due to occlusions, poor lighting, and irregular structures. Although transformer-based methods achieve high accu...
U-CFR: Uncertainty-Guided Cascade Forward Refinement for Interactive Segmentation
arXiv:2607. 20705v1 Announce Type: cross Abstract: Interactive image segmentation is critical for efficient image annotation; however, existing methods often require many corrective clicks or rely on passive refinement schemes that converge slowly.