arXiv Computer Vision By Sheekar Banerjee, Md. Srabon Chowdhury, Md. Mahbub Hasan Akash, Ishtiak Al Mamoon

Lightweight Vision Transformer-Based U-Net for Brain Tumor Segmentation from MRI

Read the original on arXiv Computer Vision →

The paper introduces a lightweight Vision Transformer‑based U‑Net for brain tumor segmentation from MRI, combining U‑Net’s hierarchical feature extraction with a compact ViT bottleneck to capture both local and global context. With only 2.6 million trainable parameters, the model achieves a mean Intersection over Union of 0.8100 and a Dice score of 0.8446 on the TCGA LGG dataset, surpassing the baseline U‑Net by 3.75% and 3.15% respectively. Extensive quantitative and qualitative analyses, including confusion matrices, precision‑recall curves, and tumor size dependency studies, demonstrate the method’s effectiveness and robustness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Sep 16

NeuroTS-Net: Multi-Class Semantic Segmentation of Pediatric Brain Tumors in Multi-Modal MRI

NeuroTS-Net is a 3‑D encoder‑decoder CNN designed for multi‑class semantic segmentation of pediatric brain tumors in multi‑modal MRI. It uses a dual‑scale raw‑detail stream, adaptive low‑resolution context selection, and detail‑preserving multipath downsampling to maintain fine intensity and boundary information while modeling broader tumor context. Trained on the BraTS 2026 pediatric dataset, it outperformed nnU‑Net and MedNeXt, achieving Dice scores of 0.938/0.937 on internal validation and 0.927/0.926 on the official challenge set.

By Darius Peteleaza, Razvan-Gabriel Dumitru, Bogdan Neamtu, Arpad Gellert, Mariana Sandu, Claudiu Matei
arXiv Computer Vision
Sep 22

VGG16-MCA UNet: Whole-Tumor Segmentation in 2D FLAIR MRI with Decoder-Side Channel Attention

VGG16-MCA UNet is a hybrid neural network that combines an ImageNet‑pretrained VGG16 encoder with a decoder enhanced by a Multi‑Channel Attention module, trained using Focal Tversky loss to address class imbalance. The model was evaluated as a 2‑D, FLAIR‑only whole‑tumor segmenter on BraTS 2020 and LGG datasets, achieving a pixel‑level Dice of 95.10 % on BraTS and 88.32 % on LGG in a 5‑fold cross‑validation setting. Inference time is 66.32 ms per 256×256 slice on a single RTX 2060, only slightly slower than a VGG16‑UNet without attention. whyItMatters":"The study provides a reproducible 2‑D FLAIR baseline for whole‑tumor segmentation, demonstrating high Dice scores and detailed reporting of training and evaluation protocols."

By Shubham Gajjar, Deep Joshi, Avi Poptani, Vishal Barot