arXiv:2608.30371v1 Announce Type: new
Abstract: Automatic cardiac image segmentation is pivotal for diagnosing and treating cardiac diseases. In this work, we introduce MCSeg, a volumetric transforme...
By Zhiyu Ye, Hairong Zheng, Tong Zhang
OptiModNet is a lightweight UNet‑Transformer hybrid designed for optic disc and cup segmentation. It incorporates grouped‑query and channel attention across multiple stages, along with an Aggregated Pyramid Loss to improve gradient flow and structural consistency. Evaluated on the REFUGE2 dataset, it surpasses existing methods by over 2.5 % while using only 3.73 GFLOPs and 1.93 M parameters.
By Soumili Ghosh, Debapriya Roy, Aryan Das, Bikash Santra
arXiv:2608. 20229v1 Announce Type: cross Abstract: Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-specific artifacts.
By Mosharof Hossain, Md Rabiul Islam, Limon Halder, Erchin Serpedin, Md Kamrul Hasan
arXiv:2608. 08575v1 Announce Type: cross Abstract: Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail.
By Zheyang Jing, Qin Lu, Jianwang Li, Yujie Yang, Chen Yi, Shaofeng Jiang
arXiv:2608.30183v1 Announce Type: cross
Abstract: Lightweight channel attention mechanisms are widely used in image classification, yet their effectiveness in fine-grained visual recognition (FGVR) r...
By Yu-Sheng Liu, Yu-Chen Tung
TailProp introduces a hierarchical vision backbone that adapts propagation dynamics across visual representations using a Tail Propagation Operator (TPO). TPO combines Gaussian and Cauchy stable-process propagators—one with rapidly decaying influence and one with heavy-tailed influence—by predicting a content-conditioned, channel-wise coefficient that fuses the two responses in the DCT domain. The resulting architecture achieves state‑of‑the‑art performance on ImageNet‑1K, Mask R‑CNN, and ADE20K, outperforming matched propagation baselines across multiple vision tasks.
TailProp introduces a hierarchical vision backbone that uses a Tail Propagation Operator (TPO) combining Gaussian and Cauchy stable-process propagators to adaptively mix rapid and heavy-tailed spatial influences. The operator predicts a content-conditioned, channel-wise coefficient that fuses the two responses in the DCT domain, achieving efficient $O(N^{1.5})$ spatial mixing. Across multiple vision tasks, TailProp consistently outperforms matched propagation baselines, achieving state‑of‑the‑art results on ImageNet, COCO, and ADE20K.
By Jiahao Kong, Zihan Li
arXiv:2609.24494v1 Announce Type: new
Abstract: Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth e...
By Xuezhi Xiang, Jiayao Liu, Heqi Xiang, Yuqi Hu, Yiming Chen, Shanjun Zhang
arXiv:2609.22506v1 Announce Type: new
Abstract: Vision Transformers allocate most parameters to multi-layer perceptrons (MLPs) for channel mixing, while token interactions usually rely on quadratic m...
By Ali Mehizel, Oussama Khaldi
arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu
arXiv:2608.30975v1 Announce Type: cross
Abstract: Cardiac magnetic resonance imaging (CMR) produces rich sequential data such as temporal cine videos and spatial LGE/mapping stacks, yet most deep lea...
By Athira J. Jacob, Puneet Sharma, Dorin Comaniciu, Daniel Rueckert
arXiv:2608. 08713v1 Announce Type: cross Abstract: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges.
By Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein