Hugging Face Trending Papers

TailProp: content-adaptive light- and heavy-tailed propagation for vision

Read the original on Hugging Face Trending Papers →

TailProp introduces a hierarchical vision backbone that adapts propagation dynamics across visual representations using a Tail Propagation Operator (TPO). TPO combines Gaussian and Cauchy stable-process propagators—one with rapidly decaying influence and one with heavy-tailed influence—by predicting a content-conditioned, channel-wise coefficient that fuses the two responses in the DCT domain. The resulting architecture achieves state‑of‑the‑art performance on ImageNet‑1K, Mask R‑CNN, and ADE20K, outperforming matched propagation baselines across multiple vision tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 11

TailProp: content-adaptive light- and heavy-tailed propagation for vision

TailProp introduces a hierarchical vision backbone that uses a Tail Propagation Operator (TPO) combining Gaussian and Cauchy stable-process propagators to adaptively mix rapid and heavy-tailed spatial influences. The operator predicts a content-conditioned, channel-wise coefficient that fuses the two responses in the DCT domain, achieving efficient $O(N^{1.5})$ spatial mixing. Across multiple vision tasks, TailProp consistently outperforms matched propagation baselines, achieving state‑of‑the‑art results on ImageNet, COCO, and ADE20K.

By Jiahao Kong, Zihan Li
arXiv Computer Vision
Sep 21

Extending Decoupled Attention to Dense Prediction and Masked Training for Multi-Channel Images

The paper introduces an extension of the Decoupled Vision Transformer (DC‑ViT) to handle multi‑channel imaging (MCI) data, where each channel carries a distinct semantic signal. By tokenizing each channel separately and then decoupling intra‑channel from inter‑channel updates, the model preserves channel‑specific features. The authors further enable independent per‑channel masking by solving a linear assignment between retained patches, allowing masked training without restricting visible tokens. Experiments on fluorescence microscopy, imaging mass cytometry, and satellite imaging datasets demonstrate that this approach outperforms the strongest Multi‑Channel Vision Transformer baselines on both classification and dense‑prediction tasks.

By Umar Marikkar, Sameed Husain, Muhammad Awais, Sara Atito
arXiv Computer Vision
Sep 22

0.5\%>100\%: Bidirectional Reciprocal Learning for Referring Image Segmentation

The paper introduces Bidirectional Reciprocal Learning (BRL), a parameter‑efficient fine‑tuning framework for referring image segmentation that operates on frozen vision foundation models. BRL employs two lightweight adapters—Reciprocal Attention Adapter (RAA) for token‑level cross‑modal attention and Reciprocal Gate Adapter (RGA) for channel‑level gating—to enable hierarchical, bidirectional information flow between vision and language. Experiments on RefCOCO, RefCOCO+, and RefCOCOg show that BRL outperforms existing methods while updating fewer than 0.5% of backbone parameters.

By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu
arXiv AI
Aug 25

Hyperbolic Hierarchical Clustering for Visual Representation Learning

The paper introduces ClusterMixer, a token mixer based on hierarchical clustering in hyperbolic space, designed to be transparent and interpretable. It forms the core of a new vision backbone called HCFormer, which incorporates multiple clustering strategies to maintain strong performance. Experiments show HCFormer surpasses existing backbones on tasks such as image classification, object detection, instance segmentation, and semantic segmentation.

By Jianan Wei, Guikun Chen, Zhiyuan Weng, Chunchao Guo, Yujia Wang, Wenguan Wang