arXiv:2512. 06208v3 Announce Type: replace-cross Abstract: Inference of standard convolutional neural networks (CNNs) on FPGAs often incurs high latency and a long initiation interval due to the deep nested loops required to densely convolve every input pixel regardless of its feature value.
By Ho Fung Tsoi, Dylan Rankin, Vladimir Loncar, Philip Harris
arXiv:2609.22807v1 Announce Type: new
Abstract: Implicit Neural Representation (INR) leverages neural networks to represent discrete signals such as images as continuous ones, where the network weigh...
By Jinglei Shi, Xinran Chang, Jiaqi Cui, Yingjie Xia, Zhaolin Xiao, Chongyi Li
arXiv:2609.19122v1 Announce Type: new
Abstract: Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit...
By Meng'en Qin, Yinchen Liu, Mingxuan Cui, Youlu Xing
arXiv:2607. 06918v1 Announce Type: cross Abstract: Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks.
By Sojung An, Junha Lee, Sujeong You, Nam Ik Cho, Donghyun Kim
The paper introduces a training‑adaptive convolutional sparse coding (CSC) framework that learns the sparsity coefficient jointly with network parameters using an unfolded FISTA optimization. By treating the coefficient as a differentiable variable, the method balances information retention and compression through an information bottleneck perspective, promoting compact yet task‑relevant representations. A label‑free post‑training strategy further adjusts compression for corrupted inputs, yielding competitive accuracy on clean data and enhanced robustness to perturbations on CIFAR and ImageNet.
By Meng'en Qin, Yinchen Liu, Mingxuan Cui, Youlu Xing
arXiv:2608.30183v1 Announce Type: cross
Abstract: Lightweight channel attention mechanisms are widely used in image classification, yet their effectiveness in fine-grained visual recognition (FGVR) r...
By Yu-Sheng Liu, Yu-Chen Tung
arXiv:2607. 03012v1 Announce Type: cross Abstract: Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation.
By Dongyeun Lee, Amir Zandieh, Vahab Mirrokni, Junmo Kim, Insu Han
arXiv:2607. 02097v1 Announce Type: cross Abstract: Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory access from gather-based computation; while Large Kernel Acceleration (LKA) helps on small feature maps, it becomes counterproductive on large feature maps, even slower than non-accelerated implementations.
By Wan Song, Wei Zhou, Rui Wang, Jun Yu, Toru Kurihara, Jiajia Xu, Shu Zhan
arXiv:2607. 11940v1 Announce Type: cross Abstract: As the scale of large pre-trained models continues to grow, fine-tuning them under limited memory budgets has become increasingly challenging.
By Gengyu Zhang, Haiyin Ran, Zhengbao He, Yuhang Liu, Hanling Tian, Zhehao Huang, Xiaolin Huang
arXiv:2606. 27802v1 Announce Type: new Abstract: Hierarchical predictive coding provides an interpretable framework for perception as error-driven inference in multi-layer generative models, while sparse coding imposes parsimonious latent representations through explicit sparsity constraints.
By Kazuhisa Fujita
The paper introduces Right In-Place (RiP) convolution, a memory‑efficient strategy that corrects and generalizes previous in‑place convolution formulations to arbitrary stride, dilation, padding, and rectangular kernels. RiP aligns each layer’s input and output within a shared workspace, enabling safe, row‑major access with minimal memory overhead. Experiments on 10,000 random layers and 84 layers from 25 architectures show no corruption, matching or improving on existing herringbone workspaces while reducing memory usage by up to 24.8% and lowering peak activation memory on Raspberry Pi Pico MCUs by 12.5–33.3% without affecting cycle counts.
By Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe
Visual Prompting (VP) has emerged as an efficient paradigm for adapting large-scale pre-trained vision models to downstream tasks by incorporating learnable prompts at the input level. However, existing VP methods typically employ dense pixel-level prompts, which often suffer from redundant perturbations, limited generalization and energy inefficiency.