arXiv:2607. 22718v1 Announce Type: cross Abstract: We propose parameter-efficient SSM-based U-Net architectures for 3D medical image segmentation.
By Mohammad Arafat Hussain, Ellen Grant, Yangming Ou
LiteViLNet is a lightweight RGB‑geometry fusion network for road segmentation that uses a MobileNetV3 RGB encoder and a tiny depth‑wise‑separable geometry encoder. Its multi‑scale fusion module enhances modality‑specific features, performs cross‑modal interaction, and applies adaptive gating, while a depth‑wise large‑kernel bridge expands contextual support with minimal overhead. The U‑Net‑style decoder is trained with deep supervision, achieving state‑of‑the‑art performance on KITTI and ORFD benchmarks and running at up to 68.73 FPS on a Jetson Orin NX with TensorRT FP16.
By Daojie Peng, Bingtao Wang, Fulong Ma, Liang Zhang, Jun Ma
Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak targets in infrared imagery usually occupy only a few pixels and are easily submerged by cloud clutter, ground edges, and bright noise, making it difficult for lightweight segmentation-based methods to preserve local target structures while suppressing background interference.
VGG16-MCA UNet is a hybrid neural network that combines an ImageNet‑pretrained VGG16 encoder with a decoder enhanced by a Multi‑Channel Attention module, trained using Focal Tversky loss to address class imbalance. The model was evaluated as a 2‑D, FLAIR‑only whole‑tumor segmenter on BraTS 2020 and LGG datasets, achieving a pixel‑level Dice of 95.10 % on BraTS and 88.32 % on LGG in a 5‑fold cross‑validation setting. Inference time is 66.32 ms per 256×256 slice on a single RTX 2060, only slightly slower than a VGG16‑UNet without attention.
whyItMatters":"The study provides a reproducible 2‑D FLAIR baseline for whole‑tumor segmentation, demonstrating high Dice scores and detailed reporting of training and evaluation protocols."
By Shubham Gajjar, Deep Joshi, Avi Poptani, Vishal Barot
arXiv:1812.00877v2 Announce Type: replace-cross
Abstract: Segmentation of skin lesion boundaries in dermoscopic imaging is an important prerequisite step for computer-aided diagnosis of malignant mel...
By Glib Kechyn
arXiv:2607. 14568v1 Announce Type: cross Abstract: A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.
By A. C. Opus, J. Q. Lu
arXiv:2609.09634v1 Announce Type: new
Abstract: Large networks and ensembles often lead medical image segmentation challenges, but their storage and inference demands complicate deployment. We presen...
By Giorgi Nikvashvili, Hanxue Gu, Jie Bao, Kang Wang, Yang Yang
arXiv:2607. 02185v1 Announce Type: cross Abstract: Deep learning has achieved remarkable performance in medical image segmentation, yet it suffers from critical limitations: mathematical intractability, substantial parameter requirements, and lack of clinical interpretability.
By Mohammad Amanour Rahman
arXiv:2608. 08575v1 Announce Type: cross Abstract: Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail.
By Zheyang Jing, Qin Lu, Jianwang Li, Yujie Yang, Chen Yi, Shaofeng Jiang
OptiModNet is a lightweight UNet‑Transformer hybrid designed for optic disc and cup segmentation. It incorporates grouped‑query and channel attention across multiple stages, along with an Aggregated Pyramid Loss to improve gradient flow and structural consistency. Evaluated on the REFUGE2 dataset, it surpasses existing methods by over 2.5 % while using only 3.73 GFLOPs and 1.93 M parameters.
By Soumili Ghosh, Debapriya Roy, Aryan Das, Bikash Santra
The paper introduces FHEAT, a fractional diffusion operator derived from the discrete cosine transform, to learn how much spectral mixing each stage of a 3D medical segmentation network should perform. By reparameterizing the operator with a semigroup time, the authors enable the optimizer to decide whether global mixing is needed, resulting in a lightweight U‑shaped architecture (Light‑UNETR) paired with a Kolmogorov‑Arnold mixer (KAN3D). In semi‑supervised experiments, FHEAT‑Seg achieves state‑of‑the‑art Dice scores while dramatically reducing FLOPs through learned spectral sparsification.
By Yi-Hui Shen, Tie-Qiang Li
The paper presents a two‑pipeline framework for retinal fundus analysis that combines four‑class disease classification with vessel segmentation. It fine‑tunes eight ImageNet‑pretrained CNNs on the FIVES dataset, applies five gradient‑based explanation methods to assess model interpretability, and benchmarks ten U‑Net variants—including transformer‑based and attention‑enhanced architectures—on the FIVES and DRIVE datasets. The best classification results come from ResNet101 (94.17% accuracy), while the strongest segmentation performance is achieved by Attention U‑Net with a ResNet101V2 backbone, improving DRIVE IoU from 60.80% to 64.83%.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Pronay Debnath, Asif Iftekher Fahim, Faisal Muhammad Shah