arXiv Computer Vision By Umar Marikkar, Sameed Husain, Muhammad Awais, Sara Atito

Extending Decoupled Attention to Dense Prediction and Masked Training for Multi-Channel Images

Read the original on arXiv Computer Vision →

The paper introduces an extension of the Decoupled Vision Transformer (DC‑ViT) to handle multi‑channel imaging (MCI) data, where each channel carries a distinct semantic signal. By tokenizing each channel separately and then decoupling intra‑channel from inter‑channel updates, the model preserves channel‑specific features. The authors further enable independent per‑channel masking by solving a linear assignment between retained patches, allowing masked training without restricting visible tokens. Experiments on fluorescence microscopy, imaging mass cytometry, and satellite imaging datasets demonstrate that this approach outperforms the strongest Multi‑Channel Vision Transformer baselines on both classification and dense‑prediction tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 20

OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation

OptiModNet is a lightweight UNet‑Transformer hybrid designed for optic disc and cup segmentation. It incorporates grouped‑query and channel attention across multiple stages, along with an Aggregated Pyramid Loss to improve gradient flow and structural consistency. Evaluated on the REFUGE2 dataset, it surpasses existing methods by over 2.5 % while using only 3.73 GFLOPs and 1.93 M parameters.

By Soumili Ghosh, Debapriya Roy, Aryan Das, Bikash Santra
Hugging Face Trending Papers
Sep 10

TailProp: content-adaptive light- and heavy-tailed propagation for vision

TailProp introduces a hierarchical vision backbone that adapts propagation dynamics across visual representations using a Tail Propagation Operator (TPO). TPO combines Gaussian and Cauchy stable-process propagators—one with rapidly decaying influence and one with heavy-tailed influence—by predicting a content-conditioned, channel-wise coefficient that fuses the two responses in the DCT domain. The resulting architecture achieves state‑of‑the‑art performance on ImageNet‑1K, Mask R‑CNN, and ADE20K, outperforming matched propagation baselines across multiple vision tasks.