CATCH is a conditional 3D diffusion model operating in an invertible Haar-wavelet domain designed to inpaint masked regions in T1‑weighted brain MRI with plausible, tumor‑free tissue while preserving observed anatomy. The model’s denoiser uses noisy target coefficients, voided‑image coefficients, and a signed mask, guided by tumor‑excluded wavelet reconstruction and a hole‑focused loss, and hard compositing ensures observed voxels remain unchanged. Experiments on BraTS data show that a weighted mixture of tumor‑derived, irregular‑blob, and ellipsoidal masks yields the best performance, achieving higher SSIM, PSNR, and lower MSE compared to fixed or random augmentation baselines.
By Simon Winther Albertsen, Hjalte Bjoernstrup, Said Djafar Said, Mostafa Mehdipour Ghazi
The paper introduces an unsupervised approach to medical image segmentation by training a Denoising Diffusion Probabilistic Model (DDPM) on 21 unlabeled abdominal CT scans to learn anatomical features. The encoder weights from the DDPM are transferred to a U‑Net for downstream segmentation on the BTCV multi‑organ dataset, resulting in a significant Dice score improvement for liver segmentation from 0.75 to 0.93. In low‑data regimes, diffusion‑pretrained models retain robust performance, achieving high Dice scores even with only 10% of labeled data.
By Akshat G, Divyansh Gupta, Shaleen Bhatnagar, Shilpa Ankalaki, Tusar Kanti Mishra
Acquiring pixel-level annotations for medical image segmentation is a severe bottleneck. Traditional U-Net architectures, while effective, learn local texture patterns and lack awareness of global ana...
The paper introduces FHEAT, a fractional diffusion operator derived from the discrete cosine transform, to learn how much spectral mixing each stage of a 3D medical segmentation network should perform. By reparameterizing the operator with a semigroup time, the authors enable the optimizer to decide whether global mixing is needed, resulting in a lightweight U‑shaped architecture (Light‑UNETR) paired with a Kolmogorov‑Arnold mixer (KAN3D). In semi‑supervised experiments, FHEAT‑Seg achieves state‑of‑the‑art Dice scores while dramatically reducing FLOPs through learned spectral sparsification.
By Yi-Hui Shen, Tie-Qiang Li
The paper presents a segmentation pipeline for brain metastases in both pre‑ and post‑treatment cases using a 5‑fold nnU‑Net ResEnc‑L ensemble trained on 1,296 four‑modality cases. A rule‑based post‑processing cascade improves the lesion‑wise Dice similarity coefficient (LW‑DSC) for enhancing tumour, tumour core, whole tumour, and resection cavity sub‑regions, achieving LW‑DSC scores of 0.733, 0.751, 0.713, and 0.549 respectively on the official validation leaderboard. The authors conduct a five‑fold out‑of‑fold analysis to validate the robustness of each post‑processing stage, provide a mechanistic explanation of LW‑DSC behaviour, and report thirteen negative results that challenge common intuitions, with all code released under Apache‑2.0.
By Haobin Liu, Xin Wang
arXiv:2608. 16377v1 Announce Type: cross Abstract: Instance-level lesion detection has been an increasingly larger focal point in medical image segmentation besides the more standard voxel-level overlap.
By Qinghui Liu, Jon Andr\'e Ottesen, Atle Bj{\o}rnerud, Kyrre Eeg Emblem
arXiv:2607. 03103v1 Announce Type: cross Abstract: Clinical cardiac imaging pipelines currently deploy separate models for each dataset and modality, incurring redundant training costs and precluding knowledge sharing across anatomically related tasks.
By Jiahao Liu, Hang Wei, Shuai Wu
arXiv:2606. 03069v1 Announce Type: cross Abstract: Generalized segmentation of medical images prevents performance degradation when different imaging devices and clinical protocols are used across multiple domains.
By Aqsa Naseer, Maryam Bibi, Syeda Samiya Urooj, Muhammad Khurram Shahzad
The paper introduces CalSAM, a lightweight adaptation framework that fine‑tunes only the mask decoder of the Segment Anything Model (SAM) while keeping its encoders frozen. CalSAM employs a Feature Fisher Information Penalty (FIP) to reduce encoder sensitivity to domain shift and a Confidence Misalignment Penalty (CMP) to curb overconfident voxel‑wise errors. Experiments on cross‑center, scanner‑shift, and motion‑corrupted brain MRI datasets show significant gains in Dice similarity coefficient, Hausdorff distance, and expected calibration error, with only a modest training‑time overhead.
By Behraj Khan, Tahir Qasim Syed, Syed Ahmad Chan Bukhari
The study investigates how few expert-annotated cases are needed to fine‑tune MedSAM3 for abdominal organ segmentation using Low‑Rank Adaptation (LoRA). With only 10 annotated CT or MRI cases, the LoRA‑adapted models achieve performance comparable to specialist systems that require orders of magnitude more data, including reliable gallbladder segmentation and near‑state‑of‑the‑art results for liver, kidneys, and spleen. The approach also generalizes to cardiac segmentation on the Whole Heart dataset, and training takes only 3–5 hours per organ on a single GPU, roughly twice as fast as nnU-Net.
By Sachin Dudda Nagaraju, Bendik Skarre Abrahamsen, Ashkan Moradi, Mattijs Elschot
arXiv:2606. 15370v1 Announce Type: cross Abstract: This work demonstrates a full reproduction and extension of MNet, a hybrid 2D/3D convolutional network designed for anisotropic medical image segmentation.
By Kirsten Odendaal, Rade Bajic
Medical image segmentation is often framed as a search for stronger architectures, but this can obscure a more fundamental question: what does the dataset require from the model? In medical imaging, this requirement is shaped by foreground occupancy, morphology, boundary ambiguity, topology sensitivity, annotation quality, acquisition variation, and operating point.