Extending Decoupled Attention to Dense Prediction and Masked Training for Multi-Channel Images
Read the original on arXiv Computer Vision →The paper introduces an extension of the Decoupled Vision Transformer (DC‑ViT) to handle multi‑channel imaging (MCI) data, where each channel carries a distinct semantic signal. By tokenizing each channel separately and then decoupling intra‑channel from inter‑channel updates, the model preserves channel‑specific features. The authors further enable independent per‑channel masking by solving a linear assignment between retained patches, allowing masked training without restricting visible tokens. Experiments on fluorescence microscopy, imaging mass cytometry, and satellite imaging datasets demonstrate that this approach outperforms the strongest Multi‑Channel Vision Transformer baselines on both classification and dense‑prediction tasks.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.