arXiv Computer Vision

MCL: Meta Convolution Layer

arXiv AI
Sep 10

S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network

The paper introduces S$^3$F-Net, a dual‑branch network that fuses spatial and spectral representations for medical image classification. It combines a deep spatial CNN with a shallow spectral encoder, SpectraNet, which uses a learnable SpectralFilter layer to process the full Fourier spectrum efficiently. Evaluated on four medical imaging datasets, S$^3$F-Net consistently outperforms spatial‑only baselines, achieving state‑of‑the‑art accuracy on BRISC2025 and surpassing deeper models on the Chest X‑Ray Pneumonia dataset.

By Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan
arXiv Machine Learning
2d ago

HAND: A Biologically-Inspired Activation Function that Improves Generalisation and Sample Efficiency in Image Classification

The paper introduces HAND, a biologically-inspired activation function that incorporates homeostasis, accelerating nonlinearity, and divisive normalization to act as an inductive bias in deep neural networks. Experiments on image classification show that using HAND allows a ConvNeXt-tiny model to reach ImageNet1k accuracy in 25 epochs versus 200 epochs for the baseline, and yields larger accuracy gains on long-tailed and reduced-data settings. The authors report that HAND does not degrade generalisation on common corruptions and can improve the model’s ability to reject unknown classes, with benefits observed across multiple CNN architectures and datasets.

By Michael W. Spratling, Heiko H. Sch\"utt
Hugging Face Trending Papers
Aug 11

Grid-Preserving Knowledge Distillation: Transferring Convolutional Inductive Bias to Vision Transformers under Data Scarcity

Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher is an effective remedy that leaves the deployed model unchanged. General-purpose feature distillation, however, transfers little in this setting.

arXiv Machine Learning
Jun 3

Dynamic Short Convolutions Improve Transformers

arXiv:2606. 03825v1 Announce Type: new Abstract: Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forward layers, residual connections, and normalization.

By Oliver Sieberling, Bharat Runwal, Rameswar Panda, Yoon Kim
arXiv Computer Vision
Sep 4

ProgResViT: Progressive Resolution and Width for Adaptive Vision Transformers

ProgResViT is an input‑adaptive Vision Transformer that processes images progressively across multiple rounds, starting with a low‑resolution image and a narrow subnetwork and refining the prediction with higher resolution and a wider subnetwork if needed. The method introduces Progress‑Conditioned Soft Gating (PSG) to share a single backbone across rounds while conditioning token fusion and layer outputs on the current round, block, and input resolution. Experiments on DeiT show improved accuracy‑compute trade‑offs compared to adaptive‑width, adaptive‑depth, and dynamic‑token baselines, and the design also benefits self‑supervised DINO representations and downstream semantic segmentation.

By Ali Hojjat, Janek Haberer, Olaf Landsiedel