arXiv Computer Vision

LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder

arXiv Computer Vision
Sep 17

LiteViLNet: Lightweight Vision-LiDAR Fusion Network for Efficient Road Segmentation

LiteViLNet is a lightweight RGB‑geometry fusion network for road segmentation that uses a MobileNetV3 RGB encoder and a tiny depth‑wise‑separable geometry encoder. Its multi‑scale fusion module enhances modality‑specific features, performs cross‑modal interaction, and applies adaptive gating, while a depth‑wise large‑kernel bridge expands contextual support with minimal overhead. The U‑Net‑style decoder is trained with deep supervision, achieving state‑of‑the‑art performance on KITTI and ORFD benchmarks and running at up to 68.73 FPS on a Jetson Orin NX with TensorRT FP16.

By Daojie Peng, Bingtao Wang, Fulong Ma, Liang Zhang, Jun Ma
Hugging Face Trending Papers
Jul 27

LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target Detection

Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak targets in infrared imagery usually occupy only a few pixels and are easily submerged by cloud clutter, ground edges, and bright noise, making it difficult for lightweight segmentation-based methods to preserve local target structures while suppressing background interference.

arXiv Computer Vision
Sep 22

VGG16-MCA UNet: Whole-Tumor Segmentation in 2D FLAIR MRI with Decoder-Side Channel Attention

VGG16-MCA UNet is a hybrid neural network that combines an ImageNet‑pretrained VGG16 encoder with a decoder enhanced by a Multi‑Channel Attention module, trained using Focal Tversky loss to address class imbalance. The model was evaluated as a 2‑D, FLAIR‑only whole‑tumor segmenter on BraTS 2020 and LGG datasets, achieving a pixel‑level Dice of 95.10 % on BraTS and 88.32 % on LGG in a 5‑fold cross‑validation setting. Inference time is 66.32 ms per 256×256 slice on a single RTX 2060, only slightly slower than a VGG16‑UNet without attention. whyItMatters":"The study provides a reproducible 2‑D FLAIR baseline for whole‑tumor segmentation, demonstrating high Dice scores and detailed reporting of training and evaluation protocols."

By Shubham Gajjar, Deep Joshi, Avi Poptani, Vishal Barot
arXiv AI
Aug 20

OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation

OptiModNet is a lightweight UNet‑Transformer hybrid designed for optic disc and cup segmentation. It incorporates grouped‑query and channel attention across multiple stages, along with an Aggregated Pyramid Loss to improve gradient flow and structural consistency. Evaluated on the REFUGE2 dataset, it surpasses existing methods by over 2.5 % while using only 3.73 GFLOPs and 1.93 M parameters.

By Soumili Ghosh, Debapriya Roy, Aryan Das, Bikash Santra
arXiv AI
Sep 24

Learning Spectral Allocation: A Fractional Diffusion Framework for Adaptive Volumetric Segmentation

The paper introduces FHEAT, a fractional diffusion operator derived from the discrete cosine transform, to learn how much spectral mixing each stage of a 3D medical segmentation network should perform. By reparameterizing the operator with a semigroup time, the authors enable the optimizer to decide whether global mixing is needed, resulting in a lightweight U‑shaped architecture (Light‑UNETR) paired with a Kolmogorov‑Arnold mixer (KAN3D). In semi‑supervised experiments, FHEAT‑Seg achieves state‑of‑the‑art Dice scores while dramatically reducing FLOPs through learned spectral sparsification.

By Yi-Hui Shen, Tie-Qiang Li
arXiv Computer Vision
Sep 4

Explainable Convolutional Neural Networks for Retinal Fundus Classification and Cutting-Edge Segmentation Models for Retinal Blood Vessels from Fundus Images

The paper presents a two‑pipeline framework for retinal fundus analysis that combines four‑class disease classification with vessel segmentation. It fine‑tunes eight ImageNet‑pretrained CNNs on the FIVES dataset, applies five gradient‑based explanation methods to assess model interpretability, and benchmarks ten U‑Net variants—including transformer‑based and attention‑enhanced architectures—on the FIVES and DRIVE datasets. The best classification results come from ResNet101 (94.17% accuracy), while the strongest segmentation performance is achieved by Attention U‑Net with a ResNet101V2 backbone, improving DRIVE IoU from 60.80% to 64.83%.

By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Pronay Debnath, Asif Iftekher Fahim, Faisal Muhammad Shah