STATrack is a fully spiking neural network designed for UAV visual tracking using only RGB inputs, eliminating the need for costly event cameras. It introduces Adaptive Mutual Information Maximization (AMIM) to preserve fine-grained target information in deep spiking representations and a sample-difficulty-aware dynamic weighting strategy to adjust the mutual‑information constraint during training. Experiments on four UAV tracking benchmarks show that STATrack achieves state‑of‑the‑art performance while maintaining low theoretical energy consumption.
By Pengzhi Zhong, Jiwei Mo, Dan Zeng, Feixiang He, Shuiwang Li
arXiv:2609.37047v1 Announce Type: cross
Abstract: We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key r...
By Aidin Attar, Eleonora Cicciarella, Michele Rossi
Mask IPL introduces a Computation Graph Clipping technique that removes noise from Intrinsic Position Learning (IPL) in spiking neural networks for event-based tracking. By applying a validity mask to each layer, the method aligns forward and backward propagation with ideal gradients without adding parameters. Experiments show that Mask IPL consistently improves tracking performance on Tiny‑scale and Base‑scale benchmarks such as FE108, FELT, and VisEvent.
By Yimeng Shan, Malu Zhang
arXiv:2607. 11914v1 Announce Type: cross Abstract: A central goal of current Spiking Neural Network (SNN) research is to improve their accuracy toward becoming low-power alternatives to Artificial Neural Networks (ANNs).
By Jiahong Zhang, Sijun Shen, Man Yao, Han Xu, Mingqiang Huang, Yonghong Tian, Bo Xu, Guoqi Li
Visual Prompting (VP) has emerged as an efficient paradigm for adapting large-scale pre-trained vision models to downstream tasks by incorporating learnable prompts at the input level. However, existing VP methods typically employ dense pixel-level prompts, which often suffer from redundant perturbations, limited generalization and energy inefficiency.
arXiv:2609.39514v1 Announce Type: new
Abstract: Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. How...
By Shuai Wang, Malu Zhang, Mingquan Liu, Weihui Dai, Dehao Zhang, Jieyuan Zhang, Yimeng Shan, Zijian Zhou, Yang Yang