TailProp introduces a hierarchical vision backbone that adapts propagation dynamics across visual representations using a Tail Propagation Operator (TPO). TPO combines Gaussian and Cauchy stable-process propagators—one with rapidly decaying influence and one with heavy-tailed influence—by predicting a content-conditioned, channel-wise coefficient that fuses the two responses in the DCT domain. The resulting architecture achieves state‑of‑the‑art performance on ImageNet‑1K, Mask R‑CNN, and ADE20K, outperforming matched propagation baselines across multiple vision tasks.
arXiv:2606. 07647v1 Announce Type: cross Abstract: Large vision language models (LVLMs) have made rapid advancements and are deployed across various applications, yet hallucinations remain a major challenge.
By Ruipeng Zhang, Zhihao Li, C. L. Philip Chen, Tong Zhang
The paper introduces an extension of the Decoupled Vision Transformer (DC‑ViT) to handle multi‑channel imaging (MCI) data, where each channel carries a distinct semantic signal. By tokenizing each channel separately and then decoupling intra‑channel from inter‑channel updates, the model preserves channel‑specific features. The authors further enable independent per‑channel masking by solving a linear assignment between retained patches, allowing masked training without restricting visible tokens. Experiments on fluorescence microscopy, imaging mass cytometry, and satellite imaging datasets demonstrate that this approach outperforms the strongest Multi‑Channel Vision Transformer baselines on both classification and dense‑prediction tasks.
By Umar Marikkar, Sameed Husain, Muhammad Awais, Sara Atito
arXiv:2609.24510v2 Announce Type: replace
Abstract: Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs...
By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu
The paper introduces Bidirectional Reciprocal Learning (BRL), a parameter‑efficient fine‑tuning framework for referring image segmentation that operates on frozen vision foundation models. BRL employs two lightweight adapters—Reciprocal Attention Adapter (RAA) for token‑level cross‑modal attention and Reciprocal Gate Adapter (RGA) for channel‑level gating—to enable hierarchical, bidirectional information flow between vision and language. Experiments on RefCOCO, RefCOCO+, and RefCOCOg show that BRL outperforms existing methods while updating fewer than 0.5% of backbone parameters.
By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu
The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.
By Soshun Kihara, Shunsuke Yasuki, Masato Taki